Agent skill

System And Data Design

by AnastasiyaW in AnastasiyaW/codex-claude-code-config

Decide whether the system will hold, and where the data lives: requirements and load first, then back-of-the-envelope numbers, building blocks (cache, queue, load balancer, CDN), and the data layer…

MITAuto-check passedDatabases

Install System And Data Design

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill system-and-data-design -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config system-and-data-design --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/development/system-and-data-design .claude/skills/system-and-data-design && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
system-and-data-design
GitHub stars
154
Token cost
~1.7k tokens
SKILL.md length
773 words
Files
16 (incl. references)
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Decide whether the system will hold, and where the data lives: requirements and load first, then back-of-the-envelope numbers, building blocks (cache, queue, load balancer, CDN), and the data layer…

  • Works in 4 steps: requirements before architecture → back-of-the-envelope, before any diagram → building blocks, each with a reason → …
  • Scaling anything
  • SKILL.md covers Scope guard — read first, Step 1 — requirements before…, Step 2 — back-of-the-envelope,… and Step 3 — building blocks, each…, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

System And Data Design is an agent skill from AnastasiyaW/codex-claude-code-config. Decide whether the system will hold, and where the data lives: requirements and load first, then back-of-the-envelope numbers, building blocks (cache, queue, load balancer, CDN), and the data layer in depth — storage engines, indexes, replication, partitioning, transactions and consistency, batch vs stream. Use when sizing or scaling anything; choosing a database, cache, queue or index; when asked "will this hold", "how many machines", "which database", "do we need a queue", "read replica", "sharding", "eventual…

Its SKILL.md is about 1.7k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including reference files (for example `references/ddia-systems-original.md`, `references/ddia-systems/batch-stream.md` and `references/ddia-systems/data-models.md`).

It sits in Databases, covering Database administration, Cloud networking and Refactoring. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • Scaling anything
  • Choosing a database
  • Asked will this hold
  • How many machines

Example prompts

  • “will this hold”
  • “how many machines”
  • “which database”
  • “/system-and-data-design”

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. requirements before architecture
  2. back-of-the-envelope, before any diagram
  3. building blocks, each with a reason
  4. the data layer

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

System And Data Design loads about 1.7k tokens when it runs, and up to ~50k if it reads all its reference files. Until then it costs about 256 tokens; SKILL.md has 773 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~256
When it runs · the whole SKILL.md, loaded when a task matches
~1.7k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~50k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 773 words, ~1,681 tokens.

Download SKILL.mdSave it as .claude/skills/system-and-data-design/SKILL.md (or your agent's skills folder). This skill also uses 15 other files; get the full folder from GitHub.
name
system-and-data-design
description
Decide whether the system will hold, and where the data lives: requirements and load first, then back-of-the-envelope numbers, building blocks (cache, queue, load balancer, CDN), and the data layer in depth — storage engines, indexes, replication, partitioning, transactions and consistency, batch vs stream. Use when sizing or scaling anything; choosing a database, cache, queue or index; when asked "will this hold", "how many machines", "which database", "do we need a queue", "read replica", "sharding", "eventual consistency", "why is this query slow at scale"; when designing an ingestion or processing pipeline; or when a service is slow under load rather than wrong. Do NOT use for module layout, dependency direction or domain boundaries (use architecture-first), for function- and naming-level quality (use code-complexity), for restructuring code that is already too large (use refactoring-safely), or for a low-traffic internal tool where the honest answer is one process and one database.

System and data design — will it hold, and where does the data live

Two questions that are usually asked together and answered separately, badly. Capacity without storage internals gives a diagram that cannot be built; storage internals without capacity gives a database choice with no reason behind it.

Scope guard — read first

The most common failure here is answering at the wrong scale.

SituationHonest answer
Internal tool, tens of usersOne process, one database, no cache. Stop.
Product with real traffic, single regionEstimate first; add a cache and a queue only where the numbers say
Multi-region, or data outgrowing one machineFull pass: estimation → building blocks → replication/partitioning → consistency

Adding a cache before measuring is the canonical way to convert one problem into two.

Step 1 — requirements before architecture

  • Functional: what must it actually do? Write it as verbs, not components.
  • Non-functional, with numbers: users, requests/sec at peak, payload size, growth, read:write ratio, acceptable latency, acceptable staleness, retention.
  • Constraints: budget, team size, existing stack, compliance, where data may live.

"Acceptable staleness" is the single most useful number and the one nobody asks for. It decides caching, replication and consistency all at once.

Step 2 — back-of-the-envelope, before any diagram

Estimate storage/day, bandwidth at peak, QPS, and working-set size. The point is not precision; it is discovering that the answer is "one machine" or "this cannot work as described" before drawing anything.

Anchors worth remembering: memory reads are ~100ns, SSD ~100µs, a same-region round trip ~0.5ms, cross-continent ~150ms. Anything crossing a network is ~1000× a memory access — which is why one N+1 query pattern outweighs most micro-optimisation.

Step 3 — building blocks, each with a reason

BlockAdd it whenCost you accept
CacheRead-heavy, tolerable staleness, measured hot setInvalidation becomes your problem
QueueWork can be async; spikes must be absorbedOrdering, retries, duplicate delivery
Read replicaReads dominate; stale reads acceptableReplication lag becomes visible to users
PartitioningOne machine cannot hold data or throughputCross-partition queries and transactions get hard
CDNStatic or cacheable content, geographically spreadPurge and versioning discipline

Each row is a trade, not an upgrade. A block added without its reason written down is a future mystery.

Step 4 — the data layer

  • Storage engines: log-structured (LSM) favours writes and compaction; B-tree favours predictable reads and in-place updates. Choose by the workload's read:write shape, not by brand.
  • Indexes are a write-cost you pay for a read-benefit. An index nobody's query plan uses is pure cost.
  • Replication: single-leader is the default and enough for most systems; multi-leader and leaderless buy availability and pay in conflict resolution you must then design.
  • Partitioning: choose the key by access pattern, not by what looks even. Hot keys and unbounded partitions are the two failures; both are visible in a histogram before they are visible in production.
  • Transactions: know which isolation level you actually get. Read-committed does not prevent lost updates; "we use transactions" is not a consistency argument.
  • Batch vs stream: batch for correctness and reprocessing, stream for freshness. Streams that cannot be replayed lose the ability to fix a bug retroactively.
Show full SKILL.md (260 more words)Show less

Review pass

  • Is every component justified by a number, or by habit?
  • What happens at 10× — which part breaks first, and is that acceptable?
  • Where can it lose data, and is that written down?
  • What is the failure mode of each dependency: degrade, queue, or fail loudly?
  • Does any user-visible read cross a replication lag nobody bounded?

References — load on demand

  • references/system-design/four-step-process.md, estimation-numbers.md, building-blocks.md, database-scaling.md, common-designs.md, reliability-operations.md
  • references/ddia-systems/storage-engines.md, data-models.md, replication.md, partitioning.md, transactions.md, batch-stream.md, fault-tolerance.md
  • *-original.md — the source skills' own framework prose, kept verbatim

Gotchas

  • Designing for a scale you do not have. The cost is paid now, the benefit maybe never. Estimate first; the estimate frequently says "one machine".
  • Cache as a fix for a slow query. It hides the query and adds invalidation. Fix the query, then decide about the cache.
  • Eventual consistency chosen by accident. A read replica added for speed silently makes some reads stale. Decide which reads may be stale, and say so.
  • The queue that became a database. Unbounded retention plus a consumer that never catches up is a data store with none of the guarantees of one.

Troubleshooting

SymptomLikely causeWhere to look
Fine in test, slow in productionWorking set exceeds memory; disk seeks per requestStorage engine, index coverage
Latency spikes at intervalsCompaction, GC, or a cron competing for IOStorage engine internals, host metrics
One shard hot, others idlePartition key follows structure, not accessPartitioning
Users see their own write disappearRead served by a lagging replicaReplication, read-your-writes
Duplicate side effectsAt-least-once delivery without idempotencyQueue semantics

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 15 other files (references) in skills/development/system-and-data-design of AnastasiyaW/codex-claude-code-config.

  • SKILL.md
  • references/ddia-systems-original.md
  • references/ddia-systems/batch-stream.md
  • references/ddia-systems/data-models.md
  • references/ddia-systems/fault-tolerance.md
  • references/ddia-systems/partitioning.md
  • references/ddia-systems/replication.md
  • references/ddia-systems/storage-engines.md
  • references/ddia-systems/transactions.md
  • references/system-design-original.md
  • references/system-design/building-blocks.md
  • references/system-design/common-designs.md
  • references/system-design/database-scaling.md
  • references/system-design/estimation-numbers.md
  • references/system-design/four-step-process.md
  • references/system-design/reliability-operations.md

Open the folder on GitHubat commit 67709af

Compare with similar skills

System And Data Design next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

System And Data Design compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
System And Data Design this skillAnastasiyaW/codex-claude-code-config154—~1.7kAutomated safety check: PassMIT
Azure Resource Manager Redis Dotnetmicrosoft/skills3.1k5 repos~3kAutomated safety check: PassMIT
Aliyun Swas Managecinience/alicloud-skills397—~1.9kAutomated safety check: PassMIT
Rustpgdogdev/pgdog5.6k—~1.9kAutomated safety check: NotesAGPL-3.0
YBA Read-Only Query CLIyugabyte/yugabyte-db11k—~1.1kAutomated safety check: PassCustom licence
Mail Timeveliovgroup/mail-time143—~1kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Official

    Azure Resource Manager SDK for Redis in .NET. An agent skill from microsoft/skills.

    3.1k GitHub starsUsed in 5 repos~3k tokens
    DatabasesAuto-check passed
  • Aliyun Swas Manage

    cinience/alicloud-skills

    A skill your agent uses when managing Alibaba Cloud Simple Application Server (SWAS OpenAPI 2020-06-01) resources end-to-end, including querying instances, starting/stopping/rebooting, executing…

    397 GitHub stars~1.9k tokensUpdated 2 mo ago
    DevOps & CloudAuto-check passed
  • Rust

    pgdogdev/pgdog

    Rust coding best practices for idiomatic, efficient, and maintainable code.

    5.6k GitHub stars~1.9k tokensUpdated today
    DatabasesAuto-check: notes
  • YBA Read-Only Query CLI

    yugabyte/yugabyte-db

    Runs read-only lookups - list, describe, get - against a live YugabyteDB Anywhere control plane using the yba CLI, with mutating commands explicitly off-limits.

    11k GitHub stars~1.1k tokensUpdated today
    DatabasesAuto-check passed
  • Mail Time

    veliovgroup/mail-time

    A skill your agent uses when building, wiring, reviewing, or debugging MailTime and ostrio:mailer email queues for horizontally scaled Node.js, Bun, or Meteor apps.

    143 GitHub stars~1k tokensUpdated 3 days ago
    DatabasesAuto-check passed
  • DB Ops Sop

    OpenDCAI/DataMind

    Database operations runbook — backup, recovery, performance tuning, troubleshooting.

    451 GitHub stars~388 tokensUpdated 21 days ago
    DatabasesAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Questions about System And Data Design

What does System And Data Design do?

Decide whether the system will hold, and where the data lives: requirements and load first, then back-of-the-envelope numbers, building blocks (cache, queue, load balancer, CDN), and the data layer…. System And Data Design is an agent skill from AnastasiyaW/codex-claude-code-config. Decide whether the system will hold, and where the data lives: requirements and load first, then back-of-the-envelope numbers, building blocks (cache, queue, load balancer, CDN), and the data layer in depth — storage engines, indexes, replication, partitioning, transactions and consistency, batch vs stream.

When should I use System And Data Design?

System And Data Design fits situations like: scaling anything; choosing a database; asked will this hold; how many machines.

How do I install System And Data Design in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill system-and-data-design -a claude-code`. Or copy the skill folder (skills/development/system-and-data-design in AnastasiyaW/codex-claude-code-config) into .claude/skills/system-and-data-design in your project. Claude Code loads it when a task matches its description.

How do I install System And Data Design in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill system-and-data-design -a codex`. Or copy the skill folder (skills/development/system-and-data-design in AnastasiyaW/codex-claude-code-config) into .agents/skills/system-and-data-design in your project. Codex loads it when a task matches its description.

Can I use System And Data Design in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill system-and-data-design -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/system-and-data-design, .gemini/skills/system-and-data-design, .github/skills/system-and-data-design and .opencode/skills/system-and-data-design in your project.

What does System And Data Design need to run?

SKILL.md names no scripts, command-line tools or credentials: System And Data Design is instructions for the agent only.

Does System And Data Design access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is System And Data Design safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does System And Data Design use?

System And Data Design is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does System And Data Design use?

About 1.7k tokens (SKILL.md is roughly 6.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 49k tokens, read only when the agent opens those files.

What are the alternatives to System And Data Design?

Skills that share tags, products or a category with System And Data Design: Azure Resource Manager Redis Dotnet (microsoft/skills, 3.1k stars), Aliyun Swas Manage (cinience/alicloud-skills, 397 stars), Rust (pgdogdev/pgdog, 5.6k stars) and YBA Read-Only Query CLI (yugabyte/yugabyte-db, 11k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains System And Data Design?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.