Agent skill

Model Deployment

by secondsky in secondsky/claude-skills

Deploy ML models with FastAPI, Docker, Kubernetes. An agent skill from secondsky/claude-skills.

MITAuto-check passedDevOps & Cloud

Install Model Deployment

skills CLI
$ npx skills add secondsky/claude-skills --skill model-deployment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install secondsky/claude-skills model-deployment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/secondsky/claude-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugins/model-deployment/skills/model-deployment .claude/skills/model-deployment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-deployment
GitHub stars
227
Token cost
~2.4k tokens
SKILL.md length
447 words
Files
5 (incl. references)
Skills in repo
169
Repo updated
First seen
Licence
MIT

At a glance

Deploy ML models with FastAPI, Docker, Kubernetes. An agent skill from secondsky/claude-skills.

  • Works in 8 steps: No Health Checks = Downtime → Model Not Found Errors in Container → Unhandled Input Validation = 500 Errors → …
  • Serving predictions
  • SKILL.md covers Deployment Options, FastAPI Model Server, Docker Deployment and Model Monitoring, plus 5 more sections
  • Calls kubectl and docker

What it does

Model Deployment is an agent skill from secondsky/claude-skills. Deploy ML models with FastAPI, Docker, Kubernetes. Use for serving predictions, containerization, monitoring, drift detection, or encountering latency issues, health check failures, version conflicts.

Its SKILL.md is about 2.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including reference files (for example `references/cicd-ml-models.md`, `references/containerization-deployment.md` and `references/fastapi-production-server.md`).

It sits in DevOps & Cloud, covering Containers, Machine learning and Backend development. It works with Docker, FastAPI and Kubernetes. The repository describes itself as: Production-ready skills for Claude Code CLI - Cloudflare, React, Tailwind v4, and AI integrations. The licence is MIT.

When your agent uses it

  • Serving predictions
  • Containerization
  • Drift detection
  • Encountering latency issues

Example prompts

  • “/model-deployment”

Requirements

  • Python 3
  • Docker

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. No Health Checks = Downtime
  2. Model Not Found Errors in Container
  3. Unhandled Input Validation = 500 Errors
  4. No Drift Monitoring = Silent Degradation
  5. Missing Resource Limits = OOM Kills
  6. No Rollback Plan = Stuck on Bad Deploy
  7. Synchronous Prediction = Slow Batch Processing
  8. No CI/CD Validation = Deploy Bad Models

What it can do on your machine

Read from SKILL.md and the folder at commit 8837836. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • kubectl
    • docker

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use kubectl and docker, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Deployment loads about 2.4k tokens when it runs, and up to ~16k if it reads all its reference files. Until then it costs about 54 tokens; SKILL.md has 447 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~54
When it runs · the whole SKILL.md, loaded when a task matches
~2.4k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~16k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from secondsky/claude-skills at commit 8837836, republished under its MIT licence (© secondsky). 447 words, ~2,395 tokens.

Download SKILL.mdSave it as .claude/skills/model-deployment/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
model-deployment
description
Deploy ML models with FastAPI, Docker, Kubernetes. Use for serving predictions, containerization, monitoring, drift detection, or encountering latency issues, health check failures, version conflicts.
license
MIT
metadata.keywords
model deployment, FastAPI, Docker, Kubernetes, ML serving, model monitoring, drift detection, A/B testing, CI/CD, mlops, production ml, model versioning…

ML Model Deployment

Deploy trained models to production with proper serving and monitoring.

Deployment Options

MethodUse CaseLatency
REST APIWeb servicesMedium
BatchLarge-scale processingN/A
StreamingReal-timeLow
EdgeOn-deviceVery low

FastAPI Model Server

python
from fastapi import FastAPI
from pydantic import BaseModel
import joblib
import numpy as np

app = FastAPI()
model = joblib.load('model.pkl')

class PredictionRequest(BaseModel):
    features: list[float]

class PredictionResponse(BaseModel):
    prediction: float
    probability: float

@app.get('/health')
def health():
    return {'status': 'healthy'}

@app.post('/predict', response_model=PredictionResponse)
def predict(request: PredictionRequest):
    features = np.array(request.features).reshape(1, -1)
    prediction = model.predict(features)[0]
    probability = model.predict_proba(features)[0].max()
    return PredictionResponse(prediction=prediction, probability=probability)

Docker Deployment

dockerfile
FROM python:3.12-slim

WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

COPY model.pkl .
COPY app.py .

EXPOSE 8000
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]

Model Monitoring

python
class ModelMonitor:
    def __init__(self):
        self.predictions = []
        self.latencies = []

    def log_prediction(self, input_data, prediction, latency):
        self.predictions.append({
            'input': input_data,
            'prediction': prediction,
            'latency': latency,
            'timestamp': datetime.now()
        })

    def detect_drift(self, reference_distribution):
        # Compare current predictions to reference
        pass

Deployment Checklist

  • Model validated on test set
  • API endpoints documented
  • Health check endpoint
  • Authentication configured
  • Logging and monitoring setup
  • Model versioning in place
  • Rollback procedure documented

Quick Start: Deploy Model in 6 Steps

bash
# 1. Save trained model
import joblib
joblib.dump(model, 'model.pkl')

# 2. Create FastAPI app (see references/fastapi-production-server.md)
# app.py with /predict and /health endpoints

# 3. Create Dockerfile
cat > Dockerfile << 'EOF'
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY app.py model.pkl ./
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]
EOF

# 4. Build and test locally
docker build -t model-api:v1.0.0 .
docker run -p 8000:8000 model-api:v1.0.0

# 5. Push to registry
docker tag model-api:v1.0.0 registry.example.com/model-api:v1.0.0
docker push registry.example.com/model-api:v1.0.0

# 6. Deploy to Kubernetes
kubectl apply -f deployment.yaml
kubectl rollout status deployment/model-api

Known Issues Prevention

1. No Health Checks = Downtime

Problem: Load balancer sends traffic to unhealthy pods, causing 503 errors.

Solution: Implement both liveness and readiness probes:

python
# app.py
@app.get("/health")  # Liveness: Is service alive?
async def health():
    return {"status": "healthy"}

@app.get("/ready")  # Readiness: Can handle traffic?
async def ready():
    try:
        _ = model_store.model  # Verify model loaded
        return {"status": "ready"}
    except:
        raise HTTPException(503, "Not ready")
yaml
# deployment.yaml
livenessProbe:
  httpGet:
    path: /health
    port: 8000
  initialDelaySeconds: 30
readinessProbe:
  httpGet:
    path: /ready
    port: 8000
  initialDelaySeconds: 5
2. Model Not Found Errors in Container

Problem: FileNotFoundError: model.pkl when container starts.

Solution: Verify model file is copied in Dockerfile and path matches:

dockerfile
# ❌ Wrong: Model in wrong directory
COPY model.pkl /app/models/  # But code expects /app/model.pkl

# ✅ Correct: Consistent paths
COPY model.pkl /models/model.pkl
ENV MODEL_PATH=/models/model.pkl

# In Python:
model_path = os.getenv("MODEL_PATH", "/models/model.pkl")
3. Unhandled Input Validation = 500 Errors

Problem: Invalid inputs crash API with unhandled exceptions.

Solution: Use Pydantic for automatic validation:

python
from pydantic import BaseModel, Field, validator

class PredictionRequest(BaseModel):
    features: List[float] = Field(..., min_items=1, max_items=100)

    @validator('features')
    def validate_finite(cls, v):
        if not all(np.isfinite(val) for val in v):
            raise ValueError("All features must be finite")
        return v

# FastAPI auto-validates and returns 422 for invalid requests
@app.post("/predict")
async def predict(request: PredictionRequest):
    # Request is guaranteed valid here
    pass
4. No Drift Monitoring = Silent Degradation

Problem: Model performance degrades over time, no one notices until users complain.

Solution: Implement drift detection (see references/model-monitoring-drift.md):

python
monitor = ModelMonitor(reference_data=training_data, drift_threshold=0.1)

@app.post("/predict")
async def predict(request: PredictionRequest):
    prediction = model.predict(features)
    monitor.log_prediction(features, prediction, latency)

    # Alert if drift detected
    if monitor.should_retrain():
        alert_manager.send_alert("Model drift detected - retrain recommended")

    return prediction
5. Missing Resource Limits = OOM Kills

Problem: Pod killed by Kubernetes OOMKiller, service goes down.

Solution: Set memory/CPU limits and requests:

yaml
resources:
  requests:
    memory: "512Mi"  # Guaranteed
    cpu: "500m"
  limits:
    memory: "1Gi"    # Max allowed
    cpu: "1000m"

# Monitor actual usage:
kubectl top pods
6. No Rollback Plan = Stuck on Bad Deploy

Problem: New model version has bugs, no way to revert quickly.

Solution: Tag images with versions, keep previous deployment:

bash
# Deploy with version tag
kubectl set image deployment/model-api model-api=registry/model-api:v1.2.0

# If issues, rollback to previous
kubectl rollout undo deployment/model-api

# Or specify version
kubectl set image deployment/model-api model-api=registry/model-api:v1.1.0
7. Synchronous Prediction = Slow Batch Processing

Problem: Processing 10,000 predictions one-by-one takes hours.

Solution: Implement batch endpoint:

python
@app.post("/predict/batch")
async def predict_batch(request: BatchPredictionRequest):
    # Process all at once (vectorized)
    features = np.array(request.instances)
    predictions = model.predict(features)  # Much faster!
    return {"predictions": predictions.tolist()}
8. No CI/CD Validation = Deploy Bad Models

Problem: Deploying model that fails basic tests, breaking production.

Solution: Validate in CI pipeline (see references/cicd-ml-models.md):

yaml
# .github/workflows/deploy.yml
- name: Validate model performance
  run: |
    python scripts/validate_model.py \
      --model model.pkl \
      --test-data test.csv \
      --min-accuracy 0.85  # Fail if below threshold
Show full SKILL.md (190 more words)Show less

Best Practices

  • Version everything: Models (semantic versioning), Docker images, deployments
  • Monitor continuously: Latency, error rate, drift, resource usage
  • Test before deploy: Unit tests, integration tests, performance benchmarks
  • Deploy gradually: Canary (10%), then full rollout
  • Plan for rollback: Keep previous version, document procedure
  • Log predictions: Enable debugging and drift detection
  • Set resource limits: Prevent OOM kills and resource contention
  • Use health checks: Enable proper load balancing

When to Load References

Load reference files for detailed implementations:

  • FastAPI Production Server: Load references/fastapi-production-server.md for complete production-ready FastAPI implementation with error handling, validation (Pydantic models), logging, health/readiness probes, batch predictions, model versioning, middleware, exception handlers, and performance optimizations (caching, async)

  • Model Monitoring & Drift: Load references/model-monitoring-drift.md for ModelMonitor implementation with KS-test drift detection, Jensen-Shannon divergence, Prometheus metrics integration, alert configuration (Slack, email), continuous monitoring service, and dashboard endpoints

  • Containerization & Deployment: Load references/containerization-deployment.md for multi-stage Dockerfiles, model versioning in containers, Docker Compose setup, A/B testing with Nginx, Kubernetes deployments (rolling update, blue-green, canary), GitHub Actions CI/CD, and deployment checklists

  • CI/CD for ML Models: Load references/cicd-ml-models.md for complete GitHub Actions pipeline with model validation, data validation, automated testing, security scanning, performance benchmarks, automated rollback, and deployment strategies

© secondsky, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files (references) in plugins/model-deployment/skills/model-deployment of secondsky/claude-skills.

  • SKILL.md
  • references/cicd-ml-models.md
  • references/containerization-deployment.md
  • references/fastapi-production-server.md
  • references/model-monitoring-drift.md

Open the folder on GitHubat commit 8837836

Compare with similar skills

Model Deployment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Deployment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Deployment this skillsecondsky/claude-skills227—~2.4kAutomated safety check: PassMIT
Devops Excellencemajiayu000/spellbook286—~2.4kAutomated safety check: NotesMIT
Devops EngineerYikai-Liao/symusic1891 repos~1.5kAutomated safety check: PassMIT
LangBot Deployment Guidelangbot-app/LangBot18k—~1.2kAutomated safety check: NotesApache-2.0
Devopsnicepkg/auto-company1922 repos~814Automated safety check: PassMIT
Debug Openshell ClusterNVIDIA/OpenShell15k—~19kAutomated safety check: NotesApache-2.0

Similar skills

  • Devops Excellence

    majiayu000/spellbook

    DevOps and CI/CD expert. An agent skill from majiayu000/spellbook.

    286 GitHub stars~2.4k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes
  • Devops Engineer

    Yikai-Liao/symusic

    Creates Dockerfiles, configures CI/CD pipelines, writes Kubernetes manifests, and generates Terraform/Pulumi infrastructure templates.

    189 GitHub starsUsed in 1 repo~1.5k tokens
    DevOps & CloudAuto-check passed
  • LangBot Deployment Guide

    langbot-app/LangBot

    Deploys and configures a LangBot instance with Docker Compose or Kubernetes, covering config.yaml, the Box sandbox runtime, the plugin runtime and the global API key.

    18k GitHub stars~1.2k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Devops

    nicepkg/auto-company

    Deploy to Cloudflare (Workers, R2, D1), Docker, GCP (Cloud Run, GKE), Kubernetes (kubectl, Helm).

    192 GitHub starsUsed in 2 repos~814 tokens
    DevOps & CloudAuto-check passed
  • Debug Openshell Cluster

    NVIDIA/OpenShell

    Official

    Debug why an OpenShell gateway deployment is unhealthy, unreachable, or unable to create sandboxes.

    15k GitHub stars~19k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • AI Server

    Opentrons/opentrons

    Conventions for the opentrons-ai-server FastAPI service — project structure, uv dependency management, settings, testing, Docker, and deployment.

    521 GitHub stars~2.5k tokensUpdated yesterday
    DevOps & CloudAuto-check: notes

More from secondsky/claude-skills

All 169 skills in this repo
  • Tanstack AI

    secondsky/claude-skills

    TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.

    227 GitHub starsUsed in 1 repo~3.6k tokens
    Auto-check: notes
  • Auto Animate

    secondsky/claude-skills

    AutoAnimate (@formkit/auto-animate) zero-config animations for React.

    227 GitHub stars~2.9k tokensUpdated 9 days ago
    Auto-check passed
  • Base UI React

    secondsky/claude-skills

    MUI Base UI unstyled React components with Floating UI. An agent skill from secondsky/claude-skills.

    227 GitHub stars~1.9k tokensUpdated 9 days ago
    Auto-check passed
  • Cloudflare Images

    secondsky/claude-skills

    This skill should be used when the user asks to "upload images to Cloudflare", "implement direct creator upload", "configure image transformations", "optimize WebP/AVIF", "create image variants"…

    227 GitHub stars~3.6k tokensUpdated 9 days ago
    Auto-check: notes
  • Cloudflare Nextjs

    secondsky/claude-skills

    Deploy Next.js to Cloudflare Workers via the OpenNext adapter (@opennextjs/cloudflare).

    227 GitHub stars~5.3k tokensUpdated 9 days ago
    Auto-check: notes
  • Cloudflare Sandbox

    secondsky/claude-skills

    Cloudflare Sandboxes SDK for secure code execution in Linux containers at edge.

    227 GitHub stars~4.5k tokensUpdated 9 days ago
    Auto-check passed

Categories

Questions about Model Deployment

What does Model Deployment do?

Deploy ML models with FastAPI, Docker, Kubernetes. An agent skill from secondsky/claude-skills. Model Deployment is an agent skill from secondsky/claude-skills. Deploy ML models with FastAPI, Docker, Kubernetes.

When should I use Model Deployment?

Model Deployment fits situations like: serving predictions; containerization; drift detection; encountering latency issues.

How do I install Model Deployment in Claude Code?

Run `npx skills add secondsky/claude-skills --skill model-deployment -a claude-code`. Or copy the skill folder (plugins/model-deployment/skills/model-deployment in secondsky/claude-skills) into .claude/skills/model-deployment in your project. Claude Code loads it when a task matches its description.

How do I install Model Deployment in Codex?

Run `npx skills add secondsky/claude-skills --skill model-deployment -a codex`. Or copy the skill folder (plugins/model-deployment/skills/model-deployment in secondsky/claude-skills) into .agents/skills/model-deployment in your project. Codex loads it when a task matches its description.

Can I use Model Deployment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add secondsky/claude-skills --skill model-deployment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-deployment, .gemini/skills/model-deployment, .github/skills/model-deployment and .opencode/skills/model-deployment in your project.

What does Model Deployment need to run?

Going by SKILL.md and its folder, Model Deployment needs the command-line tools its instructions call (kubectl and docker). Our summary lists: Python 3; Docker.

Does Model Deployment access the network?

SKILL.md contains no URLs. Its commands use docker, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Model Deployment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Deployment use?

Model Deployment is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Deployment use?

About 2.4k tokens (SKILL.md is roughly 9.6k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 14k tokens, read only when the agent opens those files.

What are the alternatives to Model Deployment?

Skills that share tags, products or a category with Model Deployment: Devops Excellence (majiayu000/spellbook, 286 stars), Devops Engineer (Yikai-Liao/symusic, 189 stars), LangBot Deployment Guide (langbot-app/LangBot, 18k stars) and Devops (nicepkg/auto-company, 192 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Deployment?

secondsky (a GitHub user) maintains it in secondsky/claude-skills, which has 227 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 28, 2026.

Source: secondsky/claude-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.