System Design
ninehills/skills
Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues.
When designing distributed systems for scalability, reliability, and consistency.
$ npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install ancoleman/ai-design-components designing-distributed-systems --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/designing-distributed-systems .claude/skills/designing-distributed-systems && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "designing-distributed-systems" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/designing-distributed-systems into .claude/skills/designing-distributed-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "designing-distributed-systems", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/ancoleman/ai-design-components/tree/main/skills/designing-distributed-systemsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install ancoleman/ai-design-components designing-distributed-systems --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/designing-distributed-systems .agents/skills/designing-distributed-systems && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "designing-distributed-systems" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/designing-distributed-systems into .agents/skills/designing-distributed-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "designing-distributed-systems", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install ancoleman/ai-design-components designing-distributed-systems --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/designing-distributed-systems .cursor/skills/designing-distributed-systems && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "designing-distributed-systems" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/designing-distributed-systems into .cursor/skills/designing-distributed-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "designing-distributed-systems", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/ancoleman/ai-design-components.git --path skills/designing-distributed-systems--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install ancoleman/ai-design-components designing-distributed-systems --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/designing-distributed-systems .gemini/skills/designing-distributed-systems && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "designing-distributed-systems" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/designing-distributed-systems into .gemini/skills/designing-distributed-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "designing-distributed-systems", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install ancoleman/ai-design-components designing-distributed-systemsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/designing-distributed-systems .github/skills/designing-distributed-systems && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "designing-distributed-systems" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/designing-distributed-systems into .github/skills/designing-distributed-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "designing-distributed-systems", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install ancoleman/ai-design-components designing-distributed-systems --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/ancoleman/ai-design-components.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/designing-distributed-systems .opencode/skills/designing-distributed-systems && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "designing-distributed-systems" agent skill from https://github.com/ancoleman/ai-design-components/tree/main/skills/designing-distributed-systems into .opencode/skills/designing-distributed-systems/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "designing-distributed-systems", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
designing-distributed-systemsWhen designing distributed systems for scalability, reliability, and consistency.
Designing Distributed Systems is an agent skill from ancoleman/ai-design-components. When designing distributed systems for scalability, reliability, and consistency. Covers CAP/PACELC theorems, consistency models (strong, eventual, causal), replication patterns (leader-follower, multi-leader, leaderless), partitioning strategies (hash, range, geographic), transaction patterns (saga, event sourcing, CQRS), resilience patterns (circuit breaker, bulkhead), service discovery, and caching strategies for building fault-tolerant distributed architectures.
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 30 other files, including reference files (for example `examples/circuit-breaker/circuit_breaker.py`, `examples/consistent-hashing/consistent_hash.py` and `examples/cqrs/cqrs_example.py`).
It sits in Backend & APIs, covering Event-driven systems, Microservices and Caching. The repository describes itself as: Comprehensive UI/UX and Backend component design skills for AI-assisted development with Claude. The licence is MIT.
Read from SKILL.md and the folder at commit 76551b7. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships script files (Python, from the files we listed), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Designing Distributed Systems loads about 4.3k tokens when it runs, and up to ~32k if it reads all its reference files. Until then it costs about 125 tokens; SKILL.md has 1,429 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from ancoleman/ai-design-components at commit 76551b7, republished under its MIT licence (© ancoleman). 1,429 words, ~4,282 tokens.
.claude/skills/designing-distributed-systems/SKILL.md (or your agent's skills folder). This skill also uses 22 other files; get the full folder from GitHub.Design scalable, reliable, and fault-tolerant distributed systems using proven patterns and consistency models.
Distributed systems are the foundation of modern cloud-native applications. Understanding fundamental trade-offs (CAP theorem, PACELC), consistency models, replication patterns, and resilience strategies is essential for building systems that scale globally while maintaining correctness and availability.
Apply when:
CAP Theorem: In a distributed system experiencing a network partition, choose between Consistency (C) or Availability (A). Partition tolerance (P) is mandatory.
Network partitions WILL occur → Always design for P
During partition:
├─ CP (Consistency + Partition Tolerance)
│ Use when: Financial transactions, inventory, seat booking
│ Trade-off: System unavailable during partition
│ Examples: HBase, MongoDB (default), etcd
│
└─ AP (Availability + Partition Tolerance)
Use when: Social media, caching, analytics, shopping carts
Trade-off: Stale reads possible, conflicts need resolution
Examples: Cassandra, DynamoDB, RiakPACELC: Extends CAP to consider normal operations (no partition).
Strong Consistency ◄─────────────────────► Eventual Consistency
│ │ │
Linearizable Causal Consistency Convergent
(Slowest, (Middle Ground, (Fastest,
Most Consistent) Causally Ordered) Eventually Consistent)Strong Consistency (Linearizability):
Eventual Consistency:
Causal Consistency:
Bounded Staleness:
1. Leader-Follower (Single-Leader):
2. Multi-Leader:
3. Leaderless (Dynamo-style):
Hash Partitioning (Consistent Hashing):
Range Partitioning:
Geographic Partitioning:
Circuit Breaker:
[Closed] → Normal operation
│ (failures exceed threshold)
▼
[Open] → Fail fast (don't call failing service)
│ (timeout expires)
▼
[Half-Open] → Try single request
│ success → [Closed]
│ failure → [Open]Bulkhead Isolation:
Timeout and Retry:
Rate Limiting and Backpressure:
Saga Pattern:
Choreography: Services react to events
Order Service → OrderCreated event
Payment Service → listens → PaymentProcessed event
Inventory Service → listens → InventoryReserved event
(Compensating: if payment fails → InventoryReleased event)Orchestration: Central coordinator
Saga Orchestrator:
1. Call Order Service
2. Call Payment Service
3. Call Inventory Service
(If step fails → call compensating transactions in reverse)Event Sourcing:
CQRS (Command Query Responsibility Segregation):
Client-Side Discovery:
Server-Side Discovery:
Service Mesh:
Cache-Aside (Lazy Loading):
Read:
1. Check cache → hit? return
2. Miss? Query database
3. Store in cache, returnWrite-Through:
Write:
1. Write to cache
2. Cache writes to database synchronously
3. Return successWrite-Behind (Write-Back):
Write:
1. Write to cache
2. Return success
3. Cache writes to database asynchronously (batched)Cache Invalidation:
Decision Tree:
├─ Money involved? → Strong Consistency
├─ Double-booking unacceptable? → Strong Consistency
├─ Causality important (chat, edits)? → Causal Consistency
├─ Read-heavy, stale tolerable? → Eventual Consistency
└─ Default? → Eventual (then strengthen if needed)├─ Single region writes? → Leader-Follower
├─ Multi-region writes + conflicts OK? → Multi-Leader
├─ Multi-region writes + no conflicts? → Leader-Follower with failover
└─ Maximum availability? → Leaderless (quorum)├─ Need range scans? → Range Partitioning (risk: hot spots)
├─ Data residency requirements? → Geographic Partitioning
└─ Default? → Hash Partitioning (consistent hashing)| System | If Partition | Else (Normal) | Use Case |
|---|---|---|---|
| Spanner | PC | EC (strong) | Global SQL |
| DynamoDB | PA | EL (eventual) | High availability |
| Cassandra | PA | EL (tunable) | Wide-column store |
| MongoDB | PC | EC (default) | Document store |
| Cosmos DB | PA/PC | EL/EC (5 levels) | Multi-model |
| Use Case | Consistency Model |
|---|---|
| Bank account balance | Strong (Linearizable) |
| Seat booking (airline) | Strong (Linearizable) |
| Inventory stock count | Strong or Bounded |
| Shopping cart | Eventual |
| Product catalog | Eventual |
| Collaborative editing | Causal |
| Chat messages | Causal |
| Social media likes | Eventual |
| DNS records | Eventual |
| Configuration | W | R | N | Consistency | Use Case |
|---|---|---|---|---|---|
| Strong | 3 | 3 | 5 | Strong | Banking |
| Balanced | 3 | 2 | 5 | Strong | Default |
| Write-heavy | 2 | 3 | 5 | Strong | Logs |
| Read-heavy | 3 | 1 | 5 | Eventual | Cache |
| Max Avail | 1 | 1 | 5 | Eventual | Analytics |
For comprehensive coverage of specific topics, see:
Complete, runnable examples demonstrating patterns:
Visual representations for complex concepts:
Related Skills:
For Kubernetes deployment: See kubernetes-operations skill for pod anti-affinity, service mesh
For infrastructure: See infrastructure-as-code skill for deploying distributed systems
For databases: See databases-sql and databases-nosql for replication configuration
For messaging: See message-queues skill for event-driven architectures, saga orchestration
For monitoring: See observability skill for distributed tracing, monitoring patterns
For testing: See performance-engineering skill for load testing distributed systems
For security: See security-hardening skill for mTLS, service authentication
1. Choose replication: Multi-leader or Leaderless
2. Partition data geographically
3. Implement conflict resolution (LWW, vector clocks, app-specific)
4. Monitor replication lag
5. Add circuit breakers between datacenters1. Define saga steps and compensating actions
2. Choose choreography (events) or orchestration (coordinator)
3. Implement idempotent handlers (retries safe)
4. Publish events with outbox pattern (transactional)
5. Monitor saga progress and timeouts1. Use leaderless replication (N=5, W=3, R=2)
2. Partition with consistent hashing
3. Add circuit breakers for failing nodes
4. Implement read repair and anti-entropy
5. Monitor quorum healthDesign for Failure:
Choose Consistency Carefully:
Idempotency is Critical:
Monitor and Observe:
Partition Strategically:
Version Everything:
Distributed Monolith:
Two-Phase Commit (2PC) Overuse:
Ignoring Network Failures:
Strong Consistency Everywhere:
No Conflict Resolution Strategy:
Cache Stampede:
Replication Lag Too High:
Split-Brain Scenario:
Hot Partitions:
Saga Timeout/Stalled:
Conflict Resolution Failures:
© ancoleman, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 22 other files (references) in skills/designing-distributed-systems of ancoleman/ai-design-components.
Open the folder on GitHubat commit 76551b7
Designing Distributed Systems next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Designing Distributed Systems this skillancoleman/ai-design-components | 526 | — | ~4.3k | Automated safety check: Pass | MIT | |
| System Designninehills/skills | 280 | — | ~4.7k | Automated safety check: Pass | MIT | |
| System Designwondelai/skills | 2.4k | — | ~4k | Automated safety check: Pass | MIT | |
| Azure Resource Manager Redis Dotnetmicrosoft/skills | 3.1k | 5 repos | ~3k | Automated safety check: Pass | MIT | |
| Stripe Projectsfossasia/eventyay | 1.7k | 5 repos | ~2k | Automated safety check: Notes | Apache-2.0 | |
| AWS Serverless Edazxkane/aws-skills | 367 | 4 repos | ~3.2k | Automated safety check: Pass | MIT |
ninehills/skills
Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues.
wondelai/skills
Design scalable distributed systems using structured approaches for load balancing, caching, database scaling, and message queues.
microsoft/skills
Azure Resource Manager SDK for Redis in .NET. An agent skill from microsoft/skills.
fossasia/eventyay
A skill your agent uses when the user wants to provision infrastructure or third-party services using Stripe Projects.
zxkane/aws-skills
AWS serverless and event-driven architecture expert based on Well-Architected Framework.
FoundatioFx/Foundatio
A skill your agent uses when working with Foundatio infrastructure abstractions for .NET -- caching, queuing, messaging, file storage, distributed locking, or background jobs.
ancoleman/ai-design-components
Builds AI chat interfaces and conversational UI with streaming responses, context management, and multi-modal support.
ancoleman/ai-design-components
Builds form components and data collection interfaces including contact forms, registration flows, checkout processes, surveys, and settings pages.
ancoleman/ai-design-components
Builds tables and data grids for displaying tabular information, from simple HTML tables to complex enterprise data grids.
ancoleman/ai-design-components
Creates comprehensive dashboard and analytics interfaces that combine data visualization, KPI cards, real-time updates, and interactive layouts.
ancoleman/ai-design-components
Designs layout systems and responsive interfaces including grid systems, flexbox patterns, sidebar layouts, and responsive breakpoints.
ancoleman/ai-design-components
Displays chronological events and activity through timelines, activity feeds, Gantt charts, and calendar interfaces.
Categories
When designing distributed systems for scalability, reliability, and consistency. Designing Distributed Systems is an agent skill from ancoleman/ai-design-components. When designing distributed systems for scalability, reliability, and consistency.
Designing Distributed Systems fits situations like: tasks that involve Event-driven systems; tasks that involve Microservices; tasks that involve Caching.
Run `npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a claude-code`. Or copy the skill folder (skills/designing-distributed-systems in ancoleman/ai-design-components) into .claude/skills/designing-distributed-systems in your project. Claude Code loads it when a task matches its description.
Run `npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a codex`. Or copy the skill folder (skills/designing-distributed-systems in ancoleman/ai-design-components) into .agents/skills/designing-distributed-systems in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add ancoleman/ai-design-components --skill designing-distributed-systems -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/designing-distributed-systems, .gemini/skills/designing-distributed-systems, .github/skills/designing-distributed-systems and .opencode/skills/designing-distributed-systems in your project.
Going by SKILL.md and its folder, Designing Distributed Systems needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Designing Distributed Systems is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 28k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Designing Distributed Systems: System Design (ninehills/skills, 280 stars), System Design (wondelai/skills, 2.4k stars), Azure Resource Manager Redis Dotnet (microsoft/skills, 3.1k stars) and Stripe Projects (fossasia/eventyay, 1.7k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
ancoleman (a GitHub user) maintains it in ancoleman/ai-design-components, which has 526 GitHub stars. The repository holds 75 skills in this directory. The repository was last updated on December 11, 2025.
Source: ancoleman/ai-design-components on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.