DB Sculptor
EliasOulkadi/shokunin
Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…
Design data systems by understanding storage engines, replication, partitioning, transactions, and consistency models.
$ npx skills add wondelai/skills --skill ddia-systems -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install wondelai/skills ddia-systems --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/wondelai/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ddia-systems .claude/skills/ddia-systems && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ddia-systems" agent skill from https://github.com/wondelai/skills/tree/main/ddia-systems into .claude/skills/ddia-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ddia-systems", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/wondelai/skills/tree/main/ddia-systemsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add wondelai/skills --skill ddia-systems -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install wondelai/skills ddia-systems --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wondelai/skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/ddia-systems .agents/skills/ddia-systems && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ddia-systems" agent skill from https://github.com/wondelai/skills/tree/main/ddia-systems into .agents/skills/ddia-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ddia-systems", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wondelai/skills --skill ddia-systems -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install wondelai/skills ddia-systems --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wondelai/skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/ddia-systems .cursor/skills/ddia-systems && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ddia-systems" agent skill from https://github.com/wondelai/skills/tree/main/ddia-systems into .cursor/skills/ddia-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ddia-systems", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/wondelai/skills.git --path ddia-systems--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add wondelai/skills --skill ddia-systems -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install wondelai/skills ddia-systems --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wondelai/skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/ddia-systems .gemini/skills/ddia-systems && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ddia-systems" agent skill from https://github.com/wondelai/skills/tree/main/ddia-systems into .gemini/skills/ddia-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ddia-systems", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install wondelai/skills ddia-systemsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add wondelai/skills --skill ddia-systems -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/wondelai/skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/ddia-systems .github/skills/ddia-systems && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ddia-systems" agent skill from https://github.com/wondelai/skills/tree/main/ddia-systems into .github/skills/ddia-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ddia-systems", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add wondelai/skills --skill ddia-systems -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install wondelai/skills ddia-systems --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/wondelai/skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/ddia-systems .opencode/skills/ddia-systems && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ddia-systems" agent skill from https://github.com/wondelai/skills/tree/main/ddia-systems into .opencode/skills/ddia-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ddia-systems", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ddia-systemsDesign data systems by understanding storage engines, replication, partitioning, transactions, and consistency models.
Ddia Systems is an agent skill from wondelai/skills. Design data systems by understanding storage engines, replication, partitioning, transactions, and consistency models. Use when the user mentions "database choice", "which database should I use", "SQL or NoSQL", "replication lag", "partitioning strategy", "consistency vs availability", "stream processing", "ACID transactions", "eventual consistency", "my queries are slow at scale", or "data is inconsistent across replicas". Also trigger when choosing a datastore, designing data pipelines, or debugging…
Its SKILL.md is about 4.2k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including reference files (for example `references/batch-stream.md`, `references/data-models.md` and `references/fault-tolerance.md`).
It sits in Databases, covering Database administration, NoSQL databases and Data pipelines and ETL. It works with SQL. The repository describes itself as: Wondel.ai Agent Skills — Business, Marketing, UX & Coding Frameworks from Bestselling Books. 50 skills + 12 guided journeys for Claude Code, Codex, Cursor & other agentskills.io… The licence is MIT.
7 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit c172996. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
No scripts in the folder and no shell commands in SKILL.md.
From the folder's file list and the shell code blocks in SKILL.md.
Links to these hosts (documentation or services it may open):
amazon.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ddia Systems loads about 4.2k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 175 tokens; SKILL.md has 1,891 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from wondelai/skills at commit c172996, republished under its MIT licence (© wondelai). 1,891 words, ~4,196 tokens.
.claude/skills/ddia-systems/SKILL.md (or your agent's skills folder). This skill also uses 7 other files; get the full folder from GitHub.A principled approach to building reliable, scalable, and maintainable data systems. Apply these principles when choosing databases, designing schemas, architecting distributed systems, or reasoning about consistency and fault tolerance.
Data outlives code. Applications are rewritten and frameworks come and go, but data persists for decades -- prioritize the long-term correctness, durability, and evolvability of the data layer. Most applications are data-intensive, not compute-intensive: the hard problems are data volume, complexity, and rate of change, and explicit consistency/availability/latency trade-offs separate robust systems from fragile ones.
Goal: 10/10. Score a data architecture by the seven Quick Diagnostic rows below: award ~1.4 points per row answered "yes" with evidence (deliberate, documented trade-off), 0 where the answer is "no" or unknown.
Report the current score, which diagnostic rows failed, and the improvements needed to reach 10/10.
Seven domains for reasoning about data-intensive systems:
Core concept: The data model shapes how you think about the problem. Relational, document, and graph models each impose different constraints and enable different query patterns.
Why it works: Choosing the wrong data model forces application code to compensate for representational mismatch, adding accidental complexity that compounds over time.
Key insights:
Code applications:
| Context | Pattern | Example |
|---|---|---|
| User profiles with nested data | Document model for self-contained aggregates | Profile, addresses, and preferences in one MongoDB document |
| Social network connections | Graph model for relationship traversal | Neo4j Cypher: MATCH (a)-[:FOLLOWS*2]->(b) for friend-of-friend |
| Financial ledger with joins | Relational model for referential integrity | PostgreSQL foreign keys between accounts, transactions, entries |
See references/data-models.md when picking relational vs document vs graph or evaluating schema-on-read -- adds the full trade-off matrix and query-language comparisons.
Core concept: Storage engines trade off read performance against write performance. Log-structured engines (LSM trees) optimize writes; page-oriented engines (B-trees) balance reads and writes.
Key insights:
Code applications:
| Context | Pattern | Example |
|---|---|---|
| High write throughput | LSM-tree engine | Cassandra or RocksDB for time-series ingestion at 100K+ writes/sec |
| Mixed read/write OLTP | B-tree engine | PostgreSQL B-tree indexes for transactional point lookups |
| Analytical queries | Column-oriented storage | ClickHouse or Parquet for scanning billions of rows, few columns |
See references/storage-engines.md when a workload is read/write-bound or you must choose indexes -- adds write/read-path diagrams, compaction strategies, column storage, and a benchmark-driven decision procedure.
Core concept: Replication keeps copies of data on multiple machines for fault tolerance, scalability, and latency reduction. The core challenge is handling changes consistently.
Why it works: Every replication strategy trades off consistency, availability, and latency. Making the trade-off explicit prevents subtle anomalies that surface only under load or failure.
Key insights:
Code applications:
| Context | Pattern | Example |
|---|---|---|
| Read-heavy web app | Single-leader with read replicas | PostgreSQL primary + read replicas behind pgBouncer |
| Multi-region writes | Multi-leader replication | CockroachDB or Spanner with bounded staleness |
| Shopping cart availability | Leaderless with merge | DynamoDB with last-writer-wins or application-level cart merge |
See references/replication.md when choosing single/multi/leaderless or debugging stale reads -- adds lag anomalies, quorum math, conflict resolution, and CRDTs.
Core concept: Partitioning (sharding) distributes data across nodes so each handles a subset, enabling horizontal scaling beyond a single machine.
Key insights:
Code applications:
| Context | Pattern | Example |
|---|---|---|
| Time-series data | Key-range partitioning by time + source | Partition by (sensor_id, date) to avoid current-day write hotspot |
| User data at scale | Hash partitioning on user ID | Cassandra consistent hashing on user_id for even distribution |
| Celebrity/hot-key problem | Key splitting with random suffix | Append random digit to hot key, fan out reads across 10 sub-partitions |
See references/partitioning.md when sharding or fighting a hot key -- adds rebalancing strategies, request routing, and local-vs-global secondary index trade-offs.
Core concept: Transactions provide safety guarantees (ACID) that simplify application code by letting you pretend failures and concurrency don't exist -- within the transaction's scope.
Why it works: Without transactions, every piece of application code must handle partial failures, races, and concurrent modification. Transactions move that complexity into the database, handled correctly once.
Key insights:
Code applications:
| Context | Pattern | Example |
|---|---|---|
| Account balance transfer | Serializable transaction | BEGIN; UPDATE accounts ... -100 WHERE id=1; UPDATE accounts ... +100 WHERE id=2; COMMIT; |
| Inventory reservation | SELECT FOR UPDATE to prevent write skew | SELECT stock FROM items WHERE id = X FOR UPDATE before decrementing |
| Cross-service operations | Saga instead of distributed transaction | Charge card, reserve inventory; on failure, run compensating refund |
See references/transactions.md when setting isolation levels or chasing a concurrency bug -- adds per-isolation anomaly tables, write-skew examples, 2PL vs SSI, and distributed-transaction pitfalls.
Core concept: Batch processing transforms bounded datasets in bulk; stream processing transforms unbounded event streams continuously. Both compute derived data.
Why it works: Separating the system of record from derived data (caches, indexes, materialized views) lets each be optimized independently and rebuilt from source when requirements change.
Key insights:
Code applications:
| Context | Pattern | Example |
|---|---|---|
| Daily analytics pipeline | Batch processing with Spark | Read day's events from S3, aggregate, write to warehouse |
| Real-time fraud detection | Stream processing with Flink | Kafka payment events, rules over 5-second tumbling windows |
| Syncing search index | Change data capture | Debezium captures PostgreSQL WAL, Kafka feeds Elasticsearch |
| Audit trail / event replay | Event sourcing | Store OrderPlaced, OrderShipped events; rebuild state by replaying |
See references/batch-stream.md when designing a pipeline or deriving data from a system of record -- adds dataflow engines, CDC wiring, windowing, and exactly-once techniques.
Core concept: Faults are inevitable; failures are not. A reliable system continues operating correctly even when individual components fail. Design for faults, not against them.
Key insights:
Code applications:
| Context | Pattern | Example |
|---|---|---|
| Service communication | Timeouts + retries with backoff | retry(max=3, backoff=exponential(base=1s, max=30s)) with jitter |
| Leader election | Consensus algorithm (Raft/Paxos) | etcd or ZooKeeper for distributed locks and leader election |
| Graceful degradation | Circuit breaker | Resilience4j: open circuit after 50% failures in 10-second window |
See references/fault-tolerance.md when tuning timeouts/retries or adding consensus -- adds fault classification, timeout-tuning math, Raft/Paxos mechanics, and safety/liveness guarantees.
| Mistake | Why It Fails | Fix |
|---|---|---|
| Choosing a database by popularity | Engines have fundamentally different trade-offs | Match storage engine to actual read/write patterns |
| Ignoring replication lag | Stale reads, phantom reads, lost updates | Implement read-your-writes and monotonic-read guarantees |
| Distributed transactions everywhere | 2PC is slow, fragile; coordinator is a SPOF | Design single-partition operations; use sagas across services |
| Hash partitioning everything | Destroys range query ability | Key-range partitioning for time-series; composite keys for locality |
| Assuming serializable isolation | Defaults are weaker; write skew appears in production | Check the actual default; use explicit locking where needed |
| Conflating batch and stream | Wrong tool adds latency or wasted complexity | Match processing model to data boundedness and latency needs |
| Treating all faults as recoverable | Corruption and Byzantine faults need different handling | Classify faults; design a recovery strategy per class |
| Question | If No | Action |
|---|---|---|
| Can you explain why you chose this database over alternatives? | Choice was familiarity, not requirements | Evaluate data model fit, read/write ratio, consistency needs, scaling path |
| Do you know your database's default isolation level? | Latent concurrency bugs | Check docs; test for write skew and phantom reads |
| Is your replication strategy explicitly chosen? | Implicit consistency/durability assumptions | Document sync vs async, failover behavior, lag tolerance |
| Can your system handle a hot partition key? | One popular entity can down the cluster | Add key-splitting or load shedding for hot keys |
| Do you separate system of record from derived data? | Every change requires migrating everything | Introduce CDC or event sourcing to decouple |
| Are timeouts and retries tuned, not defaulted? | Cascading failures or needless delays | Measure p99; set timeouts above p99, below cascade threshold |
| Have you tested failover in production conditions? | Recovery plan is theoretical | Run chaos experiments: kill leaders, partition networks, fill disks |
For the complete treatment with detailed diagrams and research references:
Martin Kleppmann is a distributed-systems researcher at the University of Cambridge and a former engineer at LinkedIn and Rapportive, known for his work on CRDTs and local-first software. His book Designing Data-Intensive Applications (2017) is the definitive reference for engineers building data systems, praised for making distributed-systems concepts accessible and practical.
© wondelai, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 7 other files (references) in ddia-systems of wondelai/skills.
Open the folder on GitHubat commit c172996
Ddia Systems next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ddia Systems this skillwondelai/skills | 2.4k | — | ~4.2k | Automated safety check: Pass | MIT | |
| DB SculptorEliasOulkadi/shokunin | 114 | — | ~3.1k | Automated safety check: Notes | MIT | |
| Database Architecture InterviewerPrepLabsAI/InterviewMentor | 112 | — | ~2.4k | Automated safety check: Pass | MIT | |
| System Design Data ArchitectureHoangNguyen0403/agent-skills-standard | 572 | — | ~910 | Automated safety check: Pass | MIT | |
| Databaseaiskillstore/marketplace | 433 | 3 repos | ~1.2k | Automated safety check: Pass | None | |
| AWS Storageaws/agent-toolkit-for-aws | 2.8k | — | ~5.8k | Automated safety check: Pass | Apache-2.0 |
EliasOulkadi/shokunin
Design database schemas with Prisma/Drizzle, PostgreSQL index strategy (B-tree, GIN, GiST, BRIN, Hash), query optimization (EXPLAIN ANALYZE), migration safety (expand/contract, zero-downtime), and…
PrepLabsAI/InterviewMentor
A Principal Database Engineer interviewer. An agent skill from PrepLabsAI/InterviewMentor.
HoangNguyen0403/agent-skills-standard
Choose and scale the data layer: SQL versus NoSQL per access pattern, single data ownership, replication and read scaling, partition key choice, hot partition and celebrity key mitigation.
aiskillstore/marketplace
Database development and operations workflow covering SQL, NoSQL, database design, migrations, optimization, and data engineering.
aws/agent-toolkit-for-aws
Selects, investigates, and compares AWS object, file, and block storage services, and answers cost, performance, configuration, security, and troubleshooting questions about storage services.
rand/cc-polymath
Automatically discover database skills when working with SQL, PostgreSQL, MongoDB, Redis, database schema design, query optimization, migrations, connection pooling, ORMs, or database selection.
wondelai/skills
Navigate the technology adoption lifecycle from early adopters to mainstream market.
wondelai/skills
Apply foundational design principles: affordances, signifiers, constraints, feedback, and conceptual models.
wondelai/skills
Run a structured 5-day process to prototype, test, and validate product ideas with real users.
wondelai/skills
Design habit-forming product loops using the Hook Model (Trigger, Action, Variable Reward, Investment).
wondelai/skills
Diagnose and fix retention problems using behavior design (B=MAP).
wondelai/skills
Design products and pricing around validated willingness to pay, from Ramanujam & Tacke's "Monetizing Innovation".
Works with
Categories
Design data systems by understanding storage engines, replication, partitioning, transactions, and consistency models. Ddia Systems is an agent skill from wondelai/skills. Design data systems by understanding storage engines, replication, partitioning, transactions, and consistency models.
Ddia Systems fits situations like: the user mentions database choice; which database should I use; replication lag; partitioning strategy.
Run `npx skills add wondelai/skills --skill ddia-systems -a claude-code`. Or copy the skill folder (ddia-systems in wondelai/skills) into .claude/skills/ddia-systems in your project. Claude Code loads it when a task matches its description.
Run `npx skills add wondelai/skills --skill ddia-systems -a codex`. Or copy the skill folder (ddia-systems in wondelai/skills) into .agents/skills/ddia-systems in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add wondelai/skills --skill ddia-systems -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ddia-systems, .gemini/skills/ddia-systems, .github/skills/ddia-systems and .opencode/skills/ddia-systems in your project.
SKILL.md names no scripts, command-line tools or credentials: Ddia Systems is instructions for the agent only.
SKILL.md names 1 domain. As links in the text: amazon.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ddia Systems is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.2k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 22k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Ddia Systems: DB Sculptor (EliasOulkadi/shokunin, 114 stars), Database Architecture Interviewer (PrepLabsAI/InterviewMentor, 112 stars), System Design Data Architecture (HoangNguyen0403/agent-skills-standard, 572 stars) and Database (aiskillstore/marketplace, 433 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
wondelai (a GitHub organization) maintains it in wondelai/skills, which has 2,371 GitHub stars. The repository holds 62 skills in this directory. The repository was last updated on September 10, 2026.
Source: wondelai/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.