Agent skill

ML Engineer

by majiayu000 in majiayu000/claude-skill-registry

Expert in building scalable ML systems, from data pipelines and model training to production deployment and monitoring.

MITAuto-check passedDevOps & Cloud

Install ML Engineer

skills CLI
$ npx skills add majiayu000/claude-skill-registry --skill ml-engineer -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install majiayu000/claude-skill-registry ml-engineer --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/majiayu000/claude-skill-registry.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-ml/ml-engineer-skill .claude/skills/ml-engineer && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-engineer
GitHub stars
666
Used in
1 other repo
Token cost
~2.8k tokens
SKILL.md length
955 words
Files
2
Skills in repo
971
Repo updated
First seen
Licence
MIT

At a glance

Expert in building scalable ML systems, from data pipelines and model training to production deployment and monitoring.

  • Works in 4 steps: Decision Framework → Core Workflows → Anti-Patterns & Gotchas → …
  • Tasks that involve MLOps
  • SKILL.md covers Purpose, When to Use, 2. Decision Framework and 3. Core Workflows, plus 3 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

ML Engineer is an agent skill from majiayu000/claude-skill-registry. Expert in building scalable ML systems, from data pipelines and model training to production deployment and monitoring.

Its SKILL.md is about 2.8k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `metadata.json`).

It sits in DevOps & Cloud, covering MLOps, Data pipelines and ETL and Fine-tuning. The repository describes itself as: Searchable Claude Code skills catalog with source-linked guides and generated registry artifacts. The licence is MIT.

When your agent uses it

  • Tasks that involve MLOps
  • Tasks that involve Data pipelines and ETL
  • Tasks that involve Fine-tuning

Example prompts

  • “/ml-engineer”

Requirements

  • Python 3
  • Docker

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Decision Framework
  2. Core Workflows
  3. Anti-Patterns & Gotchas
  4. Quality Checklist

What it can do on your machine

Read from SKILL.md and the folder at commit 000116a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Engineer loads about 2.8k tokens when it runs. Until then it costs about 33 tokens; SKILL.md has 955 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~33
When it runs · the whole SKILL.md, loaded when a task matches
~2.8k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from majiayu000/claude-skill-registry at commit 000116a, republished under its MIT licence (© majiayu000). 955 words, ~2,832 tokens.

Download SKILL.mdSave it as .claude/skills/ml-engineer/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
ml-engineer
description
Expert in building scalable ML systems, from data pipelines and model training to production deployment and monitoring.

Machine Learning Engineer

Purpose

Provides MLOps and production ML engineering expertise specializing in end-to-end ML pipelines, model deployment, and infrastructure automation. Bridges data science and production engineering with robust, scalable machine learning systems.

When to Use

  • Building end-to-end ML pipelines (Data → Train → Validate → Deploy)
  • Deploying models to production (Real-time API, Batch, or Edge)
  • Implementing MLOps practices (CI/CD for ML, Experiment Tracking)
  • Optimizing model performance (Latency, Throughput, Resource usage)
  • Setting up feature stores and model registries
  • Implementing model monitoring (Drift detection, Performance tracking)
  • Scaling training workloads (Distributed training)


2. Decision Framework

Model Serving Strategy
Need to serve predictions?
│
├─ Real-time (Low Latency)?
│  │
│  ├─ High Throughput? → **Kubernetes (KServe/Seldon)**
│  ├─ Low/Medium Traffic? → **Serverless (Lambda/Cloud Run)**
│  └─ Ultra-low latency (<10ms)? → **C++/Rust Inference Server (Triton)**
│
├─ Batch Processing?
│  │
│  ├─ Large Scale? → **Spark / Ray**
│  └─ Scheduled Jobs? → **Airflow / Prefect**
│
└─ Edge / Client-side?
   │
   ├─ Mobile? → **TFLite / CoreML**
   └─ Browser? → **TensorFlow.js / ONNX Runtime Web**
Training Infrastructure
Training Environment?
│
├─ Single Node?
│  │
│  ├─ Interactive? → **JupyterHub / SageMaker Notebooks**
│  └─ Automated? → **Docker Container on VM**
│
└─ Distributed?
   │
   ├─ Data Parallelism? → **Ray Train / PyTorch DDP**
   └─ Pipeline orchestration? → **Kubeflow / Airflow / Vertex AI**
Feature Store Decision
NeedRecommendationRationale
Simple / MVPNo Feature StoreUse SQL/Parquet files. Overhead of FS is too high.
Team ConsistencyFeastOpen source, manages online/offline consistency.
Enterprise / ManagedTecton / HopsworksFull governance, lineage, managed SLA.
Cloud NativeVertex/SageMaker FSTight integration if already in that cloud ecosystem.

Red Flags → Escalate to oracle:

  • "Real-time" training requirements (online learning) without massive infrastructure budget
  • Deploying LLMs (7B+ params) on CPU-only infrastructure
  • Training on PII/PHI data without privacy-preserving techniques (Federated Learning, Differential Privacy)
  • No validation set or "ground truth" feedback loop mechanism


3. Core Workflows

Workflow 1: End-to-End Training Pipeline

Goal: Automate model training, validation, and registration using MLflow.

Steps:

  1. Setup Tracking

    python
    import mlflow
    import mlflow.sklearn
    from sklearn.ensemble import RandomForestClassifier
    from sklearn.metrics import accuracy_score, precision_score
    
    mlflow.set_tracking_uri("http://localhost:5000")
    mlflow.set_experiment("churn-prediction-prod")
  2. Training Script (train.py)

    python
    def train(max_depth, n_estimators):
        with mlflow.start_run():
            # Log params
            mlflow.log_param("max_depth", max_depth)
            mlflow.log_param("n_estimators", n_estimators)
            
            # Train
            model = RandomForestClassifier(
                max_depth=max_depth, 
                n_estimators=n_estimators,
                random_state=42
            )
            model.fit(X_train, y_train)
            
            # Evaluate
            preds = model.predict(X_test)
            acc = accuracy_score(y_test, preds)
            prec = precision_score(y_test, preds)
            
            # Log metrics
            mlflow.log_metric("accuracy", acc)
            mlflow.log_metric("precision", prec)
            
            # Log model artifact with signature
            from mlflow.models.signature import infer_signature
            signature = infer_signature(X_train, preds)
            
            mlflow.sklearn.log_model(
                model, 
                "model",
                signature=signature,
                registered_model_name="churn-model"
            )
            
            print(f"Run ID: {mlflow.active_run().info.run_id}")
    
    if __name__ == "__main__":
        train(max_depth=5, n_estimators=100)
  3. Pipeline Orchestration (Bash/Airflow)

    bash
    #!/bin/bash
    # Run training
    python train.py
    
    # Check if model passed threshold (e.g. via MLflow API)
    # If yes, transition to Staging


Workflow 3: Drift Detection (Monitoring)

Goal: Detect if production data distribution has shifted from training data.

Steps:

  1. Baseline Generation (During Training)

    python
    import evidently
    from evidently.report import Report
    from evidently.metric_preset import DataDriftPreset
    
    # Calculate baseline profile on training data
    report = Report(metrics=[DataDriftPreset()])
    report.run(reference_data=train_df, current_data=test_df)
    report.save_json("baseline_drift.json")
  2. Production Monitoring Job

    python
    # Scheduled daily job
    def check_drift():
        # Load production logs (last 24h)
        current_data = load_production_logs()
        reference_data = load_training_data()
        
        report = Report(metrics=[DataDriftPreset()])
        report.run(reference_data=reference_data, current_data=current_data)
        
        result = report.as_dict()
        dataset_drift = result['metrics'][0]['result']['dataset_drift']
        
        if dataset_drift:
            trigger_alert("Data Drift Detected!")
            trigger_retraining()


Workflow 5: RAG Pipeline with Vector Database

Goal: Build a production retrieval pipeline using Pinecone/Weaviate and LangChain.

Steps:

  1. Ingestion (Chunking & Embedding)

    python
    from langchain.text_splitter import RecursiveCharacterTextSplitter
    from langchain_openai import OpenAIEmbeddings
    from langchain_pinecone import PineconeVectorStore
    
    # Chunking
    text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=200)
    docs = text_splitter.split_documents(raw_documents)
    
    # Embedding & Indexing
    embeddings = OpenAIEmbeddings()
    vectorstore = PineconeVectorStore.from_documents(
        docs, 
        embeddings, 
        index_name="knowledge-base"
    )
  2. Retrieval & Generation

    python
    from langchain.chains import RetrievalQA
    from langchain_openai import ChatOpenAI
    
    llm = ChatOpenAI(model="gpt-4o", temperature=0)
    
    qa_chain = RetrievalQA.from_chain_type(
        llm=llm,
        chain_type="stuff",
        retriever=vectorstore.as_retriever(search_kwargs={"k": 5})
    )
    
    response = qa_chain.invoke("How do I reset my password?")
    print(response['result'])
  3. Optimization (Hybrid Search)

    • Combine Dense Retrieval (Vectors) with Sparse Retrieval (BM25/Keywords).
    • Use Reranking (Cohere/Cross-Encoder) on the top 20 results to select best 5.


5. Anti-Patterns & Gotchas

❌ Anti-Pattern 1: Training-Serving Skew

What it looks like:

  • Feature logic implemented in SQL for training, but re-implemented in Java/Python for serving.
  • "Mean imputation" value calculated on training set but not saved; serving uses a different default.

Why it fails:

  • Model behaves unpredictably in production.
  • Debugging is extremely difficult.

Correct approach:

  • Use a Feature Store or shared library for transformations.
  • Wrap preprocessing logic inside the model artifact (e.g., Scikit-Learn Pipeline, TensorFlow Transform).
Show full SKILL.md (391 more words)Show less
❌ Anti-Pattern 2: Manual Deployments

What it looks like:

  • Data Scientist emails a .pkl file to an engineer.
  • Engineer manually copies it to a server and restarts the flask app.

Why it fails:

  • No version control.
  • No reproducibility.
  • High risk of human error.

Correct approach:

  • CI/CD Pipeline: Git push triggers build → test → deploy.
  • Model Registry: Deploy specific version hash from registry.
❌ Anti-Pattern 3: Silent Failures

What it looks like:

  • Model API returns 200 OK but prediction is garbage because input data was corrupted (e.g., all Nulls).
  • Model returns default class 0 for everything.

Why it fails:

  • Application keeps running, but business value is lost.
  • Incident detected weeks later by business stakeholders.

Correct approach:

  • Input Schema Validation: Reject bad requests (Pydantic/TFX).
  • Output Monitoring: Alert if prediction distribution shifts (e.g., if model predicts "Fraud" 0% of time for 1 hour).


7. Quality Checklist

Reliability:

  • Health Checks: /health endpoint implemented (liveness/readiness).
  • Retries: Client has retry logic with exponential backoff.
  • Fallback: Default heuristic exists if model fails or times out.
  • Validation: Inputs validated against schema before inference.

Performance:

  • Latency: P99 latency meets SLA (e.g., < 100ms).
  • Throughput: System autoscales with load.
  • Batching: Inference requests batched if using GPU.
  • Image Size: Docker image optimized (slim base, multi-stage build).

Reproducibility:

  • Versioning: Code, Data, and Model versions linked.
  • Artifacts: Saved in object storage (S3/GCS), not local disk.
  • Environment: Dependencies pinned (requirements.txt / conda.yaml).

Monitoring:

  • Technical: Latency, Error Rate, CPU/Memory/GPU usage.
  • Functional: Prediction distribution, Input data drift.
  • Business: (If possible) Attribution of prediction to outcome.

Anti-Patterns

Training-Serving Skew
  • Problem: Feature logic differs between training and serving environments
  • Symptoms: Model performs well in testing but poorly in production
  • Solution: Use feature stores or embed preprocessing in model artifacts
  • Warning Signs: Different code paths for feature computation, hardcoded constants
Manual Deployment
  • Problem: Deploying models without automation or version control
  • Symptoms: No traceability, human errors, deployment failures
  • Solution: Implement CI/CD pipelines with model registry integration
  • Warning Signs: Email/file transfers of model files, manual server restarts
Silent Failures
  • Problem: Model failures go undetected
  • Symptoms: Bad predictions returned without error indication
  • Solution: Implement input validation, output monitoring, and alerting
  • Warning Signs: 200 OK responses with garbage data, no anomaly detection
Data Leakage
  • Problem: Training data contains information not available at prediction time
  • Symptoms: Unrealistically high training accuracy, poor generalization
  • Solution: Careful feature engineering and validation split review
  • Warning Signs: Features that would only be known after prediction

© majiayu000, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/ai-ml/ml-engineer-skill of majiayu000/claude-skill-registry.

  • SKILL.md
  • metadata.json

Open the folder on GitHubat commit 000116a

Used in 1 other repository

We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in majiayu000/claude-skill-registry, which our catalogue first saw on October 7, 2026.

Compare with similar skills

ML Engineer next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Engineer compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Engineer this skillmajiayu000/claude-skill-registry6661 repos~2.8kAutomated safety check: PassMIT
ML Pipeline Workflowwshobson/agents40k12 repos~1.8kAutomated safety check: PassMIT
ML Pipeline ExpertJeffallan/claude-skills12k1 repos~1.9kAutomated safety check: PassMIT
Monitor With HaolemeHaolemeApp/Haoleme157—~1.3kAutomated safety check: PassAGPL-3.0
AWS AI MLaws/agent-toolkit-for-aws2.8k—~1.7kAutomated safety check: PassApache-2.0
ML PipelineFerroxLabs/wayland608—~3.3kAutomated safety check: PassApache-2.0

Similar skills

  • ML Pipeline Workflow

    wshobson/agents

    Guides an agent through designing an MLOps pipeline that covers data preparation, training, validation and deployment, with DAG orchestration and reference guides.

    40k GitHub starsUsed in 12 repos~1.8k tokens
    DevOps & CloudAuto-check passed
  • ML Pipeline Expert

    Jeffallan/claude-skills

    Designs ML pipeline infrastructure: experiment tracking with MLflow or Weights & Biases, Kubeflow and Airflow orchestration, Feast feature stores and model validation gates.

    12k GitHub starsUsed in 1 repo~1.9k tokens
    DevOps & CloudAuto-check passed
  • Monitor With Haoleme

    HaolemeApp/Haoleme

    Selectively monitor important long-running or resource-intensive commands with Haoleme by prefixing them with hao, so status, output, and completion notifications sync to the mobile app.

    157 GitHub stars~1.3k tokensUpdated 1 mo ago
    Data & AnalyticsAuto-check passed
  • AWS AI ML

    aws/agent-toolkit-for-aws

    Official

    Selects, deploys, and customizes AI models on Amazon SageMaker.

    2.8k GitHub stars~1.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • ML Pipeline

    FerroxLabs/wayland

    ML pipeline design covering feature engineering, model training workflows, hyperparameter tuning, cross-validation, experiment tracking (MLflow, W&B), model versioning, data versioning (DVC)…

    608 GitHub stars~3.3k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Official

    Deploys SageMaker endpoints with autoscaling, CloudWatch alarms and tags on by default, using scripts for real-time, scale-to-zero and async setups.

    11k GitHub starsUsed in 1 repo~6.9k tokens
    DevOps & CloudAuto-check passed

More from majiayu000/claude-skill-registry

All 971 skills in this repo
  • Deep Research

    majiayu000/claude-skill-registry

    Multi-source deep research using firecrawl and exa MCPs. An agent skill from majiayu000/claude-skill-registry.

    666 GitHub starsUsed in 6 repos~1.1k tokens
    Auto-check passed
  • Exa Search

    majiayu000/claude-skill-registry

    Neural search via Exa MCP for web, code, and company research.

    666 GitHub starsUsed in 5 repos~856 tokens
    Auto-check passed
  • Fal AI Media

    majiayu000/claude-skill-registry

    Unified media generation via fal.ai MCP — image, video, and audio.

    666 GitHub starsUsed in 5 repos~1.7k tokens
    Auto-check passed
  • Bgpt Paper Search

    majiayu000/claude-skill-registry

    Search scientific papers and retrieve structured experimental data extracted from full-text studies via the BGPT MCP server.

    666 GitHub starsUsed in 4 repos~619 tokens
    Auto-check: notes
  • Bio Alignment Pairwise

    majiayu000/claude-skill-registry

    Perform pairwise sequence alignment using Biopython Bio.Align.PairwiseAligner.

    666 GitHub starsUsed in 4 repos~1.7k tokens
    Auto-check passed
  • Open Notebook

    majiayu000/claude-skill-registry

    Self-hosted, open-source alternative to Google NotebookLM for AI-powered research and document analysis.

    666 GitHub starsUsed in 4 repos~2.4k tokens
    Auto-check passed

Questions about ML Engineer

What does ML Engineer do?

Expert in building scalable ML systems, from data pipelines and model training to production deployment and monitoring. ML Engineer is an agent skill from majiayu000/claude-skill-registry. Expert in building scalable ML systems, from data pipelines and model training to production deployment and monitoring.

When should I use ML Engineer?

ML Engineer fits situations like: tasks that involve MLOps; tasks that involve Data pipelines and ETL; tasks that involve Fine-tuning.

How do I install ML Engineer in Claude Code?

Run `npx skills add majiayu000/claude-skill-registry --skill ml-engineer -a claude-code`. Or copy the skill folder (skills/ai-ml/ml-engineer-skill in majiayu000/claude-skill-registry) into .claude/skills/ml-engineer in your project. Claude Code loads it when a task matches its description.

How do I install ML Engineer in Codex?

Run `npx skills add majiayu000/claude-skill-registry --skill ml-engineer -a codex`. Or copy the skill folder (skills/ai-ml/ml-engineer-skill in majiayu000/claude-skill-registry) into .agents/skills/ml-engineer in your project. Codex loads it when a task matches its description.

Can I use ML Engineer in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add majiayu000/claude-skill-registry --skill ml-engineer -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-engineer, .gemini/skills/ml-engineer, .github/skills/ml-engineer and .opencode/skills/ml-engineer in your project.

What does ML Engineer need to run?

SKILL.md names no scripts, command-line tools or credentials: ML Engineer is instructions for the agent only. Our summary lists: Python 3; Docker.

Does ML Engineer access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML Engineer safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Engineer use?

ML Engineer is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Engineer use?

About 2.8k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Engineer?

Skills that share tags, products or a category with ML Engineer: ML Pipeline Workflow (wshobson/agents, 40k stars), ML Pipeline Expert (Jeffallan/claude-skills, 12k stars), Monitor With Haoleme (HaolemeApp/Haoleme, 157 stars) and AWS AI ML (aws/agent-toolkit-for-aws, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Engineer?

majiayu000 (a GitHub user) maintains it in majiayu000/claude-skill-registry, which has 666 GitHub stars. The repository holds 971 skills in this directory. The repository was last updated on October 7, 2026.

Source: majiayu000/claude-skill-registry on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.