Agent skill

Model Deployment

by seb1n in seb1n/awesome-ai-agent-skills

Deploy trained machine learning models as production-ready services using REST APIs, containers, serverless functions, and orchestration platforms.

MITAuto-check passedBackend & APIs

Install Model Deployment

skills CLI
$ npx skills add seb1n/awesome-ai-agent-skills --skill model-deployment -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install seb1n/awesome-ai-agent-skills model-deployment --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/seb1n/awesome-ai-agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ai-ml-operations/model-deployment .claude/skills/model-deployment && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
model-deployment
GitHub stars
206
Token cost
~2.1k tokens
SKILL.md length
720 words
Files
1
Skills in repo
101
Repo updated
First seen
Licence
MIT

At a glance

Deploy trained machine learning models as production-ready services using REST APIs, containers, serverless functions, and orchestration platforms.

  • Works in 6 steps: Serialize and package the model: Export… → Build the serving API: Create a REST API… → Containerize with Docker: Write a… → …
  • The user requests model deployment
  • SKILL.md covers Workflow, Supported Technologies, Usage and Examples, plus 2 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Model Deployment is an agent skill from seb1n/awesome-ai-agent-skills. Deploy trained machine learning models as production-ready services using REST APIs, containers, serverless functions, and orchestration platforms. Use when the user requests model deployment or provides relevant inputs for this workflow.

Its SKILL.md is about 2.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Backend & APIs, covering Serverless, Containers and REST APIs. It works with Docker, Kubernetes and FastAPI. The repository describes itself as: 103 ready-to-use AI agent skills for Claude Code, OpenAI Codex, Gemini CLI, Cursor, GitHub Copilot, Windsurf, and other Agent Skills-compatible tools. Complete SKILL.md… The licence is MIT.

When your agent uses it

  • The user requests model deployment
  • Provides relevant inputs for this workflow

Example prompts

  • “/model-deployment”

Requirements

  • Python 3
  • Docker

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Serialize and package the model: Export the trained model to a portable format such as ONNX, TorchScript, SavedModel, or joblib pickle…
  2. Build the serving API: Create a REST API using FastAPI or Flask that loads the model at startup and exposes prediction endpoints. Include…
  3. Containerize with Docker: Write a Dockerfile that installs dependencies from a pinned requirements.txt, copies the model artifact and…
  4. Configure orchestration and scaling: Define Kubernetes Deployment and Service manifests (or equivalent for your platform) with resource…
  5. Deploy and verify: Push the container image to a registry, apply the Kubernetes manifests or deploy to the serverless platform, and run…
  6. Monitor and iterate: Integrate with monitoring tools like Prometheus and Grafana to track request latency, error rates, throughput, and…

What it can do on your machine

Read from SKILL.md and the folder at commit 75865a5. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python, dockerfile and yaml).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Model Deployment loads about 2.1k tokens when it runs. Until then it costs about 64 tokens; SKILL.md has 720 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~64
When it runs · the whole SKILL.md, loaded when a task matches
~2.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from seb1n/awesome-ai-agent-skills at commit 75865a5, republished under its MIT licence (© seb1n). 720 words, ~2,106 tokens.

Download SKILL.mdSave it as .claude/skills/model-deployment/SKILL.md (or your agent's skills folder).
name
model-deployment
description
Deploy trained machine learning models as production-ready services using REST APIs, containers, serverless functions, and orchestration platforms. Use when the user requests model deployment or provides relevant inputs for this workflow.
license
MIT
metadata.author
AI Agent Skills
metadata.version
1.0.0

Model Deployment

This skill enables an AI agent to deploy trained machine learning models into production environments. It covers packaging models into serving APIs with FastAPI or Flask, containerizing with Docker, orchestrating with Kubernetes, and deploying to serverless platforms. The agent handles model versioning, health checks, input validation, logging, and monitoring to ensure reliable and scalable inference in production.

Workflow

  1. Serialize and package the model: Export the trained model to a portable format such as ONNX, TorchScript, SavedModel, or joblib pickle. Bundle the model artifact with its preprocessing pipeline and any required configuration files so inference is self-contained.

  2. Build the serving API: Create a REST API using FastAPI or Flask that loads the model at startup and exposes prediction endpoints. Include a health check endpoint, request/response schemas with input validation (Pydantic models), structured logging, and error handling that returns meaningful HTTP status codes.

  3. Containerize with Docker: Write a Dockerfile that installs dependencies from a pinned requirements.txt, copies the model artifact and serving code, and sets the entrypoint to the API server. Use multi-stage builds to minimize image size and avoid including training-only dependencies.

  4. Configure orchestration and scaling: Define Kubernetes Deployment and Service manifests (or equivalent for your platform) with resource requests/limits, readiness and liveness probes pointing at the health check endpoint, and a Horizontal Pod Autoscaler to scale based on CPU, memory, or custom metrics like request latency.

  5. Deploy and verify: Push the container image to a registry, apply the Kubernetes manifests or deploy to the serverless platform, and run smoke tests against the live endpoint. Validate that responses match expected outputs for a set of known inputs.

  6. Monitor and iterate: Integrate with monitoring tools like Prometheus and Grafana to track request latency, error rates, throughput, and model-specific metrics like prediction distribution drift. Set up alerts for anomalies and establish a redeployment workflow for updated model versions using blue-green or canary strategies.

Supported Technologies

  • API frameworks: FastAPI, Flask, TorchServe, TensorFlow Serving, Triton Inference Server
  • Containerization: Docker, Podman
  • Orchestration: Kubernetes, Docker Compose, AWS ECS, Google Cloud Run
  • Serverless: AWS Lambda, Google Cloud Functions, Azure Functions
  • Monitoring: Prometheus, Grafana, Datadog, AWS CloudWatch
  • Model registries: MLflow Model Registry, AWS SageMaker Model Registry, Weights & Biases

Usage

Provide the agent with a trained model artifact, its dependencies, and the target deployment environment (local Docker, Kubernetes cluster, serverless). The agent will generate all necessary serving code, container configuration, and deployment manifests, then guide you through the deployment process.

Examples

Example 1: Deploying a Model with FastAPI
python
# app.py
import joblib
import numpy as np
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, validator
from contextlib import asynccontextmanager
from typing import List

model = None

@asynccontextmanager
async def lifespan(app: FastAPI):
    global model
    model = joblib.load("model.pkl")
    yield

app = FastAPI(title="ML Model API", version="1.0.0", lifespan=lifespan)

class PredictionRequest(BaseModel):
    features: List[float]

    @validator("features")
    def validate_features(cls, v):
        if len(v) != 4:
            raise ValueError("Expected exactly 4 features")
        return v

class PredictionResponse(BaseModel):
    prediction: int
    probability: List[float]

@app.get("/health")
def health_check():
    return {"status": "healthy", "model_loaded": model is not None}

@app.post("/predict", response_model=PredictionResponse)
def predict(request: PredictionRequest):
    try:
        features = np.array(request.features).reshape(1, -1)
        prediction = int(model.predict(features)[0])
        probability = model.predict_proba(features)[0].tolist()
        return PredictionResponse(prediction=prediction, probability=probability)
    except Exception as e:
        raise HTTPException(status_code=500, detail=str(e))
Example 2: Docker + Kubernetes Deployment

Dockerfile:

dockerfile
FROM python:3.11-slim AS builder
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt

FROM python:3.11-slim
WORKDIR /app
COPY --from=builder /usr/local/lib/python3.11/site-packages /usr/local/lib/python3.11/site-packages
COPY --from=builder /usr/local/bin/uvicorn /usr/local/bin/uvicorn
COPY app.py model.pkl ./
EXPOSE 8000
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]

k8s-deployment.yaml:

yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: ml-model-api
spec:
  replicas: 3
  selector:
    matchLabels:
      app: ml-model-api
  template:
    metadata:
      labels:
        app: ml-model-api
    spec:
      containers:
        - name: api
          image: registry.example.com/ml-model-api:v1.0.0
          ports:
            - containerPort: 8000
          resources:
            requests: { cpu: "250m", memory: "512Mi" }
            limits: { cpu: "1000m", memory: "1Gi" }
          readinessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 10
            periodSeconds: 5
          livenessProbe:
            httpGet: { path: /health, port: 8000 }
            initialDelaySeconds: 15
            periodSeconds: 10
---
apiVersion: v1
kind: Service
metadata:
  name: ml-model-api
spec:
  selector:
    app: ml-model-api
  ports:
    - port: 80
      targetPort: 8000
  type: LoadBalancer
Show full SKILL.md (297 more words)Show less

Best Practices

  • Pin all dependency versions in requirements.txt and use deterministic Docker builds to guarantee reproducibility across environments.
  • Separate model artifacts from code so you can update models without rebuilding the entire container image. Use a model registry or cloud storage with versioned paths.
  • Implement input validation with Pydantic or JSON Schema to reject malformed requests before they reach the model and produce confusing errors.
  • Use readiness probes in Kubernetes to prevent traffic from reaching pods that haven't finished loading the model, which can take significant time for large models.
  • Adopt canary deployments when releasing new model versions — route a small percentage of traffic to the new version and compare metrics before full rollout.
  • Log predictions and inputs (with PII redacted) to enable debugging, auditing, and data drift detection in production.

Edge Cases

  • Large model files (> 1 GB): Avoid baking them into Docker images. Instead, download from cloud storage (S3, GCS) at startup or mount a persistent volume. Use lazy loading if the model takes a long time to initialize.
  • Cold start latency on serverless: Serverless functions may take 10-30 seconds to load large models. Mitigate with provisioned concurrency (AWS Lambda), min-instances (Cloud Run), or by using optimized formats like ONNX Runtime.
  • Inconsistent preprocessing at inference: The preprocessing pipeline used at training must exactly match what runs at inference time. Serialize the full pipeline (e.g., with scikit-learn Pipeline + joblib) rather than reimplementing transformations separately.
  • Graceful shutdown and in-flight requests: Handle SIGTERM signals to finish processing in-flight requests before shutting down. Configure Kubernetes terminationGracePeriodSeconds to allow enough time for pending requests to complete.
  • GPU vs CPU inference mismatches: Models trained on GPU may fail if deployed to CPU-only environments. Explicitly map model tensors to CPU during loading (torch.load(path, map_location="cpu")) and test inference on the target hardware before deployment.

© seb1n, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in ai-ml-operations/model-deployment of seb1n/awesome-ai-agent-skills.

Open the folder on GitHubat commit 75865a5

Compare with similar skills

Model Deployment next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Model Deployment compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Model Deployment this skillseb1n/awesome-ai-agent-skills206—~2.1kAutomated safety check: PassMIT
Model Deploymentsecondsky/claude-skills227—~2.4kAutomated safety check: PassMIT
Deepgram Deploy Integrationjeremylongshore/tons-of-skills-marketplace2.8k—~2.5kAutomated safety check: PassMIT
Containerizing Applicationsaiskillstore/marketplace430—~1.9kAutomated safety check: PassNone
Discover Infrarand/cc-polymath181—~783Automated safety check: PassMIT
GCP Cloud Rundavila7/claude-code-templates32k7 repos~1.7kAutomated safety check: PassMIT

Similar skills

  • Model Deployment

    secondsky/claude-skills

    Deploy ML models with FastAPI, Docker, Kubernetes. An agent skill from secondsky/claude-skills.

    227 GitHub stars~2.4k tokensUpdated 11 days ago
    DevOps & CloudAuto-check passed
  • Deepgram Deploy Integration

    jeremylongshore/tons-of-skills-marketplace

    Deploy Deepgram integrations to production environments. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~2.5k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Containerizing Applications

    aiskillstore/marketplace

    Containerizes applications with Docker, docker-compose, and Helm charts.

    430 GitHub stars~1.9k tokensUpdated today
    DevOps & CloudAuto-check passed
  • Discover Infra

    rand/cc-polymath

    Automatically discover cloud, infrastructure, deployment, and container skills when working with AWS, GCP, Azure, Docker, Kubernetes, Terraform, Netlify, Heroku, serverless, or IaC

    181 GitHub stars~783 tokensUpdated 7 mo ago
    DevOps & CloudAuto-check passed
  • GCP Cloud Run

    davila7/claude-code-templates

    Specialized skill for building production-ready serverless applications on GCP.

    32k GitHub starsUsed in 7 repos~1.7k tokens
    Backend & APIsAuto-check passed
  • Polylith Project Management

    DavidVujic/python-polylith

    Create a deployable Polylith project with poly create project — a lightweight pyproject.toml under projects/<name/ that references bricks for deployment as a Docker image, wheel, AWS Lambda, GCP…

    553 GitHub stars~1.5k tokensUpdated 4 days ago
    Backend & APIsAuto-check passed

More from seb1n/awesome-ai-agent-skills

All 101 skills in this repo
  • Agent Red Teaming

    seb1n/awesome-ai-agent-skills

    Plan, execute, document, and retest authorized security assessments of AI agents and multi-agent workflows using safe adversarial cases, synthetic identities, canaries, and evidence-based findings.

    206 GitHub stars~2.8k tokensUpdated 2 mo ago
    Auto-check passed
  • Eu AI Act Readiness

    seb1n/awesome-ai-agent-skills

    Build a preliminary, evidence-based EU AI Act readiness assessment across AI-system inventory, territorial scope, operator roles, prohibited-practice screening, risk classification, transparency…

    206 GitHub stars~3.3k tokensUpdated 2 mo ago
    Auto-check passed
  • Human In The Loop

    seb1n/awesome-ai-agent-skills

    Design and verify auditable human oversight, approval gates, escalation paths, and safe state transitions for AI agent workflows.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • MCP Server Building

    seb1n/awesome-ai-agent-skills

    Design, implement, harden, and verify Model Context Protocol (MCP) servers with precise tool contracts, least-privilege authorization, safe transports, structured errors, and interoperability tests.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • PDF Processing

    seb1n/awesome-ai-agent-skills

    Inspect, extract, OCR, create, merge, split, reorder, rotate, annotate, fill, redact, compress, secure, and verify PDF documents while preserving source files and visual fidelity.

    206 GitHub stars~2.5k tokensUpdated 2 mo ago
    Auto-check passed
  • Skill Supply Chain Audit

    seb1n/awesome-ai-agent-skills

    Audit agent skills, plugins, prompts, manifests, scripts, dependencies, and bundled assets for provenance, prompt-injection, permission, execution, exfiltration, persistence, and update risk.

    206 GitHub stars~2.4k tokensUpdated 2 mo ago
    Auto-check passed

Questions about Model Deployment

What does Model Deployment do?

Deploy trained machine learning models as production-ready services using REST APIs, containers, serverless functions, and orchestration platforms. Model Deployment is an agent skill from seb1n/awesome-ai-agent-skills. Deploy trained machine learning models as production-ready services using REST APIs, containers, serverless functions, and orchestration platforms.

When should I use Model Deployment?

Model Deployment fits situations like: the user requests model deployment; provides relevant inputs for this workflow.

How do I install Model Deployment in Claude Code?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill model-deployment -a claude-code`. Or copy the skill folder (ai-ml-operations/model-deployment in seb1n/awesome-ai-agent-skills) into .claude/skills/model-deployment in your project. Claude Code loads it when a task matches its description.

How do I install Model Deployment in Codex?

Run `npx skills add seb1n/awesome-ai-agent-skills --skill model-deployment -a codex`. Or copy the skill folder (ai-ml-operations/model-deployment in seb1n/awesome-ai-agent-skills) into .agents/skills/model-deployment in your project. Codex loads it when a task matches its description.

Can I use Model Deployment in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add seb1n/awesome-ai-agent-skills --skill model-deployment -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/model-deployment, .gemini/skills/model-deployment, .github/skills/model-deployment and .opencode/skills/model-deployment in your project.

What does Model Deployment need to run?

SKILL.md names no scripts, command-line tools or credentials: Model Deployment is instructions for the agent only. Our summary lists: Python 3; Docker.

Does Model Deployment access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Model Deployment safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Model Deployment use?

Model Deployment is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Model Deployment use?

About 2.1k tokens (SKILL.md is roughly 8.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Model Deployment?

Skills that share tags, products or a category with Model Deployment: Model Deployment (secondsky/claude-skills, 227 stars), Deepgram Deploy Integration (jeremylongshore/tons-of-skills-marketplace, 2.8k stars), Containerizing Applications (aiskillstore/marketplace, 430 stars) and Discover Infra (rand/cc-polymath, 181 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Model Deployment?

seb1n (a GitHub user) maintains it in seb1n/awesome-ai-agent-skills, which has 206 GitHub stars. The repository holds 101 skills in this directory. The repository was last updated on August 9, 2026.

Source: seb1n/awesome-ai-agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.