Agent skill

ML System Design Interview

by curiositech in curiositech/some_claude_skills

Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring.

MITAuto-check passedAI & LLM Engineering

Install ML System Design Interview

skills CLI
$ npx skills add curiositech/some_claude_skills --skill ml-system-design-interview -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install curiositech/some_claude_skills ml-system-design-interview --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/curiositech/some_claude_skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/ml-system-design-interview .claude/skills/ml-system-design-interview && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-system-design-interview
GitHub stars
243
Token cost
~3.4k tokens
SKILL.md length
1,503 words
Files
5 (incl. references)
Skills in repo
109
Repo updated
First seen
Licence
MIT

At a glance

Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring.

  • L6+ design rounds
  • SKILL.md covers When to Use, The 7-Stage Design Framework, 45-Minute Time Budget and Canonical Problem Set, plus 5 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • ML architecture whiteboarding

What it does

ML System Design Interview is an agent skill from curiositech/some_claude_skills. Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring. Use for L6+ design rounds, ML architecture whiteboarding, system design practice, serving tradeoff analysis. Activate on "ML system design", "ML interview", "recommendation system design", "RAG architecture", "feature store design", "model serving". NOT for coding interviews, behavioral questions, ML theory quizzes, or paper implementations.

Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `.claude-plugin/plugin.json`, `references/evaluation-metrics-guide.md` and `references/ml-design-templates.md`).

It sits in AI & LLM Engineering, covering MLOps, LLM inference and serving and Quizzes and assessments. The repository describes itself as: Claude skills that make my life easier. The licence is MIT.

When your agent uses it

  • L6+ design rounds
  • ML architecture whiteboarding
  • System design practice
  • Serving tradeoff analysis

Example prompts

  • “ML system design”
  • “ML interview”
  • “recommendation system design”
  • “/ml-system-design-interview”

Requirements

  • Pre-approved tools (allowed-tools): Read, Write, Edit

What it can do on your machine

Read from SKILL.md and the folder at commit 6713fc7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves these tools, so the agent can use them without asking each time:

    • Read
    • Write
    • Edit

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are mermaid).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML System Design Interview loads about 3.4k tokens when it runs, and up to ~17k if it reads all its reference files. Until then it costs about 127 tokens; SKILL.md has 1,503 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~127
When it runs · the whole SKILL.md, loaded when a task matches
~3.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~17k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from curiositech/some_claude_skills at commit 6713fc7, republished under its MIT licence (© curiositech). 1,503 words, ~3,446 tokens.

Download SKILL.mdSave it as .claude/skills/ml-system-design-interview/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
ml-system-design-interview
description
Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring. Use for L6+ design rounds, ML architecture whiteboarding, system design practice, serving tradeoff analysis. Activate on "ML system design", "ML interview", "recommendation system design", "RAG architecture", "feature store design", "model serving". NOT for coding interviews, behavioral questions, ML theory quizzes, or paper implementations.
allowed-tools
Read, Write, Edit
metadata.category
Career & Interview
metadata.tags
interview, ml-system-design, machine-learning, architecture

ML System Design Interview

End-to-end ML pipeline design coaching for staff+ engineers. Covers the full arc from problem definition through production monitoring -- the scope expected at L6+ interviews at top-tier ML organizations.

This skill assumes 15+ years of ML/CV/AI/NLP experience. It does not teach fundamentals. It structures the knowledge you already have into the format interviewers reward.


When to Use

Use for:

  • Practicing 45-minute ML system design rounds
  • Structuring whiteboard presentations for recommendation, ranking, RAG, fraud, perception systems
  • Analyzing serving architecture tradeoffs (batch vs online vs streaming)
  • Identifying L6+ differentiation signals (problem ownership, org constraints, data flywheels)
  • Reviewing and critiquing ML system design answers

NOT for:

  • Coding interviews (use senior-coding-interview)
  • Behavioral / leadership questions (use interview-loop-strategist)
  • ML theory or math derivations
  • Implementing models or writing training code
  • Paper reading or research review

The 7-Stage Design Framework

Every ML system design answer follows this arc. The stages are sequential but you will loop back as constraints emerge. The Mermaid diagram below is your whiteboard skeleton.

mermaid
flowchart TD
    R[1. Requirements\n- Business goal\n- Users and scale\n- Latency/throughput SLA\n- Constraints] --> M[2. Metrics\n- Offline: precision, recall, NDCG\n- Online: CTR, conversion, revenue\n- Guardrails: latency p99, fairness]
    M --> D[3. Data\n- Sources and collection\n- Labeling strategy\n- Pipeline: ETL, validation\n- Freshness and staleness]
    D --> F[4. Features\n- Engineering and transforms\n- Feature store architecture\n- Online vs offline features\n- Freshness requirements]
    F --> Mo[5. Model\n- Architecture selection\n- Training pipeline\n- Iteration strategy\n- Baseline and ablation]
    Mo --> S[6. Serving\n- Batch vs online vs streaming\n- Caching and precomputation\n- Scaling and cost\n- Canary and shadow mode]
    S --> Mon[7. Monitoring\n- Data drift detection\n- Model degradation alerts\n- A/B testing framework\n- Rollback strategy\n- Feedback loops]
    Mon -.->|Feedback loop| D
    Mon -.->|Retrain trigger| Mo
Stage Details

Stage 1 -- Requirements (5 minutes) Ask clarifying questions before designing anything. Establish: Who is the user? What is the business metric? What is the latency SLA? What scale (QPS, data volume)? What are hard constraints (cost, privacy, regulation)? An L6+ candidate owns the problem definition -- do not wait for the interviewer to hand you requirements.

Stage 2 -- Metrics (3 minutes) Define offline metrics that you can measure before deployment AND online metrics that matter to the business. Explain the gap: "NDCG improvement offline does not always translate to CTR lift online because of position bias and novelty effects." Define guardrail metrics: latency p99, fairness across user segments, cost per prediction.

Stage 3 -- Data (7 minutes) Where does training data come from? How is it labeled (human, weak supervision, implicit signals)? What is the class balance? How fresh does data need to be? What is the data pipeline (batch ETL vs streaming)? What data quality checks exist? This stage separates L6+ candidates from L5 -- junior candidates assume clean labeled data.

Stage 4 -- Features (5 minutes) What features does the model need? Which are precomputed (offline) vs computed at request time (online)? Feature store architecture: online store (low-latency lookups) vs offline store (batch training). Feature freshness: user features update daily, item features update hourly, contextual features are real-time.

Stage 5 -- Model (8 minutes) Start with a simple baseline (logistic regression, XGBoost) and explain why. Then propose the production architecture (two-tower, transformer, etc.) and justify the upgrade. Discuss training pipeline: how often, how much data, how to handle distribution shift. Iteration strategy: what experiments to run first.

Stage 6 -- Serving (8 minutes) This is where system design and ML intersect. Discuss: inference latency requirements, batch precomputation vs online inference, GPU/CPU tradeoffs, model serving framework, caching strategy, cost optimization (quantization, distillation, spot instances). Draw the serving architecture.

Stage 7 -- Monitoring (5 minutes) What happens after deployment? Data drift detection (PSI, KL divergence). Model degradation alerts (metric decay over time). A/B testing framework (sample size, duration, novelty effects). Rollback strategy (shadow mode, canary percentage). Feedback loops that improve the model over time.


45-Minute Time Budget

PhaseMinutesWhat to Cover
Requirements + Clarification5Business goal, users, scale, SLA, constraints
Metrics3Offline, online, guardrails, metric alignment
Data7Sources, labeling, pipeline, quality, freshness
Features5Engineering, store architecture, online/offline split
Model8Baseline, production arch, training, iteration
Serving8Latency, architecture, cost, deployment strategy
Monitoring5Drift, alerts, A/B testing, rollback, feedback
Q&A Buffer4Interviewer deep-dives, defend tradeoffs

If the interviewer cuts in with questions, adapt -- but cover all 7 stages even briefly. Skipping monitoring is the most common L5 mistake.


Canonical Problem Set

ProblemKey ChallengesMust-Discuss
Recommendation SystemCold start, position bias, multi-objective optimizationTwo-tower retrieval + reranking, exploration-exploitation
Search RankingQuery intent classification, relevance vs engagement, latency at scaleInverted index + embedding retrieval, L1/L2 ranking cascade
Content ModerationMulti-modal (text+image+video), adversarial evasion, precision-recall tradeoffHuman-in-the-loop, escalation tiers, appeal workflow
RAG PipelineRetrieval quality, chunk strategy, hallucination detection, evaluationEmbedding model selection, hybrid search, reranking, citation
Fraud DetectionExtreme class imbalance, adversarial adaptation, real-time requirementFeature velocity, graph features, ensemble + rules, feedback delay
Autonomous Driving PerceptionSensor fusion, safety-critical latency, long-tail distributionMulti-task architecture, simulation, OTA updates, regulatory

Serving Architecture Comparison

PatternLatencyFreshnessCostBest For
Batch predictionN/A (precomputed)Hours-staleLow compute, high storageEmail recommendations, daily reports
Online inference10-500msReal-timeHigh compute (GPU)Search ranking, fraud detection
Near-real-time1-60sMinutes-freshMediumFeed ranking, content moderation
StreamingSub-secondContinuousHigh (always-on)Fraud, anomaly detection, bidding

Detailed serving tradeoffs, framework comparisons, and cost optimization strategies are in references/serving-tradeoffs.md.


L6+ Differentiation Signals

What separates a staff+ answer from a senior answer:

1. Own the Problem Definition Do not accept the problem as stated. Ask: "What business metric are we optimizing? Is this a revenue problem or an engagement problem? What is the current solution and why is it insufficient?" L5 candidates accept "build a recommendation system." L6+ candidates ask "what are we recommending, to whom, and what does success look like?"

2. Discuss Organizational Constraints Real systems live inside organizations. Address: team size (can we maintain a custom model or should we use a managed service?), on-call burden, cross-team data dependencies, compliance requirements, migration path from legacy system.

3. Data Flywheel Strategy Show that you think about the virtuous cycle: better model -> more engagement -> more data -> better model. Discuss how to accelerate it: active learning, implicit feedback loops, exploration strategies, cold-start bootstrapping.

4. Build vs Buy Decisions Not everything should be custom. Argue for managed services where appropriate (embedding APIs, feature stores, serving platforms) and custom solutions where competitive advantage demands it. Show you understand the total cost of ownership.

5. Multi-Objective Thinking Real systems optimize multiple objectives simultaneously: relevance AND diversity, accuracy AND fairness, quality AND latency. Discuss how to handle conflicts: Pareto optimization, constrained optimization, multi-task learning, business-rule post-processing.


Show full SKILL.md (529 more words)Show less

Whiteboard Strategy

What to draw and when:

TimeDraw ThisPurpose
0-5 minRequirements box with bullet pointsAnchor the discussion, show structured thinking
5-8 minMetric table (offline vs online)Demonstrate you think beyond model accuracy
8-15 minData pipeline diagram (sources -> ETL -> store)Show you understand data engineering
15-20 minFeature architecture (offline store + online store)Demonstrate feature store knowledge
20-28 minModel architecture + serving diagramThe core system design artifact
28-36 minFull system diagram with latency annotationsConnect everything, show you can ship
36-41 minMonitoring dashboard sketch + feedback arrowsClose the loop, show production thinking

Use boxes for components, arrows for data flow, and annotate with latency/throughput numbers. The diagram should be readable by someone who walks in at minute 30.


Anti-Patterns

Model-First Thinking

Novice: Jumps to "I would use a transformer" or "Let me describe the attention mechanism" in the first 2 minutes, before understanding the problem, defining metrics, or discussing data. Spends 70% of time on model architecture and 0% on serving.

Expert: Spends the first 10 minutes on requirements, metrics, and data before mentioning any model. Names a simple baseline first (logistic regression on handcrafted features), then argues for complexity only when the baseline's limitations are clear. Allocates equal time to serving and monitoring.

Detection: Architecture diagram has a detailed model box but no data pipeline, no feature store, no serving layer, and no monitoring component. Mentions model architecture in the first sentence.

Ignoring the Data

Novice: Assumes clean, labeled data exists at scale. Says "we would train on millions of labeled examples" without discussing where labels come from, how much they cost, what the class distribution looks like, or how stale the data gets.

Expert: Asks about data sources, labeling strategy (human vs weak supervision vs implicit signals), class imbalance handling, data freshness SLA, and data quality monitoring. Discusses the cost of labeling and proposes strategies to reduce it (active learning, semi-supervised methods, synthetic data).

Detection: No discussion of data collection, labeling costs, class imbalance, data quality checks, or data freshness anywhere in the answer. The word "label" does not appear.

No Monitoring Story

Novice: Design ends at the serving layer. No mention of what happens after the model is deployed. Does not discuss how to detect degradation, how to roll back, or how to improve the model over time.

Expert: Discusses data drift detection (population stability index, feature distribution monitoring), model performance decay alerts, A/B testing framework with proper statistical rigor, canary deployment strategy, shadow mode for safe rollouts, and explicit feedback loops that flow data back into retraining.

Detection: Architecture diagram has no monitoring component. No feedback arrows from production back to training. No mention of A/B testing, canary deployment, or rollback.


Reference Files

Consult these for deep dives -- they are NOT loaded by default:

FileConsult When
references/ml-design-templates.mdWorking through a specific problem (recommendation, search, RAG, fraud, content mod, perception). Contains 6 fully worked designs with Mermaid diagrams.
references/serving-tradeoffs.mdDeep-diving on serving architecture, framework selection, caching, cost optimization, deployment strategies. Contains framework comparisons and latency targets by use case.
references/evaluation-metrics-guide.mdChoosing metrics, understanding metric alignment, designing A/B tests, evaluating generative AI. Contains metric decision trees and formulas.

© curiositech, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in .claude/skills/ml-system-design-interview of curiositech/some_claude_skills.

  • SKILL.md
  • .claude-plugin/plugin.json
  • references/evaluation-metrics-guide.md
  • references/ml-design-templates.md
  • references/serving-tradeoffs.md

Open the folder on GitHubat commit 6713fc7

Compare with similar skills

ML System Design Interview next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML System Design Interview compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML System Design Interview this skillcuriositech/some_claude_skills243—~3.4kAutomated safety check: PassMIT
AI ML V2majiayu000/claude-skill-registry6661 repos~3.4kAutomated safety check: PassMIT
SageMaker Production Defaultshuggingface/skills11k1 repos~6.9kAutomated safety check: PassApache-2.0
Pgvector Semantic Searchtimescale/pg-aiguide1.9k1 repos~3.8kAutomated safety check: PassApache-2.0
Model Garden Deploymentgoogle/skills21k—~5kAutomated safety check: PassApache-2.0
Domodomo Local AI Maintenancedarknecrocities/DomoDomo---All-in-one-Tool239—~17kAutomated safety check: PassNone

Similar skills

  • AI ML V2

    majiayu000/claude-skill-registry

    AI/ML Workflow Bundle workflow skill. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 1 repo~3.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed
  • Pgvector Semantic Search

    timescale/pg-aiguide

    A skill your agent uses for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search.

    1.9k GitHub starsUsed in 1 repo~3.8k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Deploys open models or custom weights from Model Garden to Agent Platform endpoints, checks deployment status and cleans up endpoints, confirming before any change.

    21k GitHub stars~5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Domodomo Local AI Maintenance

    darknecrocities/DomoDomo---All-in-one-Tool

    Maintain DomoDomo private local AI features, Ollama connections, browser inference, streaming UX, embeddings, RAG, memory, and agent interfaces.

    239 GitHub stars~17k tokensUpdated 2 days ago
    AI & LLM EngineeringAuto-check passed
  • Qdrant Search Quality

    github/awesome-copilot

    Official

    Diagnoses and improves Qdrant search relevance. An agent skill from github/awesome-copilot.

    40k GitHub starsUsed in 1 repo~336 tokens
    AI & LLM EngineeringAuto-check passed

More from curiositech/some_claude_skills

All 109 skills in this repo
  • Crisis Detection Intervention AI

    curiositech/some_claude_skills

    Detect crisis signals in user content using NLP, mental health sentiment analysis, and safe intervention protocols.

    243 GitHub starsUsed in 3 repos~3.8k tokens
    Auto-check passed
  • Form Validation Architect

    curiositech/some_claude_skills

    End-to-end form handling with react-hook-form, Zod schemas, validation patterns, error messaging, field arrays, and multi-step wizards.

    243 GitHub stars~3.8k tokensUpdated 1 mo ago
    Auto-check passed
  • Competitive Cartographer

    curiositech/some_claude_skills

    Strategic analyst that maps competitive landscapes, identifies white space opportunities, and provides positioning recommendations.

    243 GitHub starsUsed in 1 repo~1.3k tokens
    Auto-check passed
  • GitHub Actions Pipeline Builder

    curiositech/some_claude_skills

    Build production CI/CD pipelines with GitHub Actions. An agent skill from curiositech/some_claude_skills.

    243 GitHub stars~2.8k tokensUpdated 1 mo ago
    Auto-check: notes
  • Computer Vision Pipeline

    curiositech/some_claude_skills

    Build production computer vision pipelines for object detection, tracking, and video analysis.

    243 GitHub starsUsed in 1 repo~4k tokens
    Auto-check passed
  • Design Archivist

    curiositech/some_claude_skills

    Long-running design anthropologist that builds comprehensive visual databases from 500-1000 real-world examples, extracting color palettes, typography patterns, layout systems, and interaction…

    243 GitHub starsUsed in 1 repo~1.6k tokens
    Auto-check passed

Questions about ML System Design Interview

What does ML System Design Interview do?

Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring. ML System Design Interview is an agent skill from curiositech/some_claude_skills. Coaches end-to-end ML system design interviews covering inference pipelines, recommendation systems, RAG, feature stores, and monitoring.

When should I use ML System Design Interview?

ML System Design Interview fits situations like: L6+ design rounds; ML architecture whiteboarding; system design practice; serving tradeoff analysis.

How do I install ML System Design Interview in Claude Code?

Run `npx skills add curiositech/some_claude_skills --skill ml-system-design-interview -a claude-code`. Or copy the skill folder (.claude/skills/ml-system-design-interview in curiositech/some_claude_skills) into .claude/skills/ml-system-design-interview in your project. Claude Code loads it when a task matches its description.

How do I install ML System Design Interview in Codex?

Run `npx skills add curiositech/some_claude_skills --skill ml-system-design-interview -a codex`. Or copy the skill folder (.claude/skills/ml-system-design-interview in curiositech/some_claude_skills) into .agents/skills/ml-system-design-interview in your project. Codex loads it when a task matches its description.

Can I use ML System Design Interview in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add curiositech/some_claude_skills --skill ml-system-design-interview -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-system-design-interview, .gemini/skills/ml-system-design-interview, .github/skills/ml-system-design-interview and .opencode/skills/ml-system-design-interview in your project.

What does ML System Design Interview need to run?

SKILL.md names no scripts, command-line tools or credentials: ML System Design Interview is instructions for the agent only. Its frontmatter pre-approves these tools: Read, Write, Edit.

Does ML System Design Interview access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML System Design Interview safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML System Design Interview use?

ML System Design Interview is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML System Design Interview use?

About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to ML System Design Interview?

Skills that share tags, products or a category with ML System Design Interview: AI ML V2 (majiayu000/claude-skill-registry, 666 stars), SageMaker Production Defaults (huggingface/skills, 11k stars), Pgvector Semantic Search (timescale/pg-aiguide, 1.9k stars) and Model Garden Deployment (google/skills, 21k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML System Design Interview?

curiositech (a GitHub organization) maintains it in curiositech/some_claude_skills, which has 243 GitHub stars. The repository holds 109 skills in this directory. The repository was last updated on September 6, 2026.

Source: curiositech/some_claude_skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.