Agent skill

ML Engineering

by magnus919 in magnus919/agent-skills

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

MITAuto-check passedAI & LLM Engineering

Install ML Engineering

skills CLI
$ npx skills add magnus919/agent-skills --skill ml-engineering -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install magnus919/agent-skills ml-engineering --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ml-engineering .claude/skills/ml-engineering && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-engineering
GitHub stars
113
Token cost
~1.5k tokens
SKILL.md length
570 words
Files
15 (incl. scripts, references)
Skills in repo
115
Repo updated
First seen
Licence
MIT

At a glance

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

  • Statistical modeling and experimental design (thats the data scientist)
  • SKILL.md covers The ML Engineer's Domain, Reference Files, Templates and Scripts, plus 3 more sections
  • Runs Python scripts from its folder
  • For operating a specific inference engine (thats a tool skill such as llama-cpp

What it does

ML Engineering is an agent skill from magnus919/agent-skills. Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity, drift response, and regression triage, grounded in practical engineering patterns for production ML systems. Do not use for statistical modeling and experimental design (that's the data scientist) or for operating a specific inference engine (that's a tool skill such as llama-cpp or vllm).

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `references/evaluation-and-lineage.md`).

It sits in AI & LLM Engineering, covering Fine-tuning and LLM inference and serving. It works with llama.cpp and vLLM. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.

When your agent uses it

  • Statistical modeling and experimental design (thats the data scientist)
  • For operating a specific inference engine (thats a tool skill such as llama-cpp

Example prompts

  • “/ml-engineering”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 96fbe07. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 2 files in scripts/ (Python), which the agent can run.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Engineering loads about 1.5k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 570 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~126
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~15k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from magnus919/agent-skills at commit 96fbe07, republished under its MIT licence (© magnus919). 570 words, ~1,465 tokens.

Download SKILL.mdSave it as .claude/skills/ml-engineering/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
ml-engineering
description
Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity, drift response, and regression triage, grounded in practical engineering patterns for production ML systems. Do not use for statistical modeling and experimental design (that's the data scientist) or for operating a specific inference engine (that's a tool skill such as llama-cpp or vllm).
license
MIT
metadata.tags
ml, machine-learning, fine-tuning, training, evaluation, quantization, mlops, inference, vllm, gguf
metadata.source_repo
https://github.com/magnus919/hermes-profiles

ML Engineering Methodology

Machine learning engineering is the bridge between model research and production systems. This methodology covers the engineering disciplines needed to train, evaluate, deploy, and maintain ML models reliably.

The ML Engineer's Domain

You ownYou don't own
Model training — LoRA/QLoRA fine-tuning, full fine-tuning, distributed trainingStatistical modeling and experimental design — that's the data scientist
Model evaluation — benchmark suites, custom eval sets, regression testingCausal inference and hypothesis testing — that's the data scientist
Quantization — GGUF, GPTQ, AWQ, bitsandbytesTraining data collection and labeling — that's the data/ML ops team
Inference serving — vLLM, llama.cpp, TGI, TritonBusiness metrics and KPI definition — that's the product manager
Evaluation harness — lm-eval-harness, custom pipelinesData pipeline architecture — that's the data engineer
Model deployment — containerization, versioning, A/B testingInfrastructure provisioning — that's the platform engineer

Reference Files

ReferenceWhen to load
references/fine-tuning.mdSetting up a LoRA/QLoRA/ full fine-tuning run — data prep, hyperparameters, validation strategy
references/evaluation.mdEvaluating a model — benchmark selection, custom eval sets, regression tracking, comparison methodology
references/evaluation-and-lineage.mdMetric/configuration decisions, repeated stochastic comparisons, end-to-end lineage, temporal feature parity, drift response, and adaptation/serving tradeoffs
references/quantization-inference.mdQuantizing a model and serving it — GGUF/GPTQ/AWQ/bitsandbytes comparison, calibration data strategies, KV cache quantization, vLLM/llama.cpp/TGI/Triton architecture, production considerations
references/training-infrastructure.mdSelecting and provisioning training infrastructure — GPU selection, VRAM budgeting, multi-GPU strategies (DDP/FSDP/DeepSpeed), cloud vs on-prem, storage, monitoring

Templates

TemplateWhen to Use
templates/training-run-record.mdRecording a training or fine-tuning run — model and data versions, full config, environment, eval results — so it can be reproduced
templates/eval-regression-table.mdTracking model quality across runs and triaging a regression — one row per eval case or capability subset
templates/quantization-decision-record.mdRecording a quantization decision — baseline, candidates compared, quality threshold, and rollback path
templates/model-lineage-record.mdLinking data/features, code/configuration, runs, artifacts, evaluations, registry state, and deployed serving versions
templates/drift-response-record.mdRecording drift signals, thresholds, diagnosis, retrain/rollback decisions, and post-action evidence

Scripts

ScriptWhen to Use
scripts/check-eval-overlap.pyChecking a training corpus against an eval corpus for test-set leakage (shared n-grams); --json for CI, exit 1 when an eval file exceeds the overlap threshold

Evals

evals/evals.json — output-quality eval manifest for this skill: fine-tuning plan review, eval-set design, quantization decision, deployment plan, regression triage, and training-run reproducibility.

Show full SKILL.md (221 more words)Show less

Core Principles

Measure before you optimize — Never quantize, prune, or distill a model without first measuring its baseline performance. Optimization without measurement is guessing.

Reproducibility is non-negotiable — Every training run needs a reproducible config: seed, data version, hyperparameters, and evaluation methodology. If you can't reproduce it, you can't ship it.

Baseline first — Before running an expensive fine-tuning run, establish a baseline with the base model. If the base model is already good enough, the fine-tuning budget is better spent elsewhere.

Test at the boundary — Model evaluation is most informative at the edges of the capability distribution, not at the center. Hard examples reveal more than easy ones.

The evaluation set is a liability — Every example in your eval set is a potential test-set leak. Use held-out sets, rotate examples, and periodically audit for contamination with the overlap checker.

When not to use

Do not use this skill for statistical modeling, experimental design, or causal inference — that's the data scientist's discipline. Do not use it to operate a specific inference engine: for llama.cpp installation, model loading, benchmarking, and troubleshooting, load the llama-cpp tool skill instead; for vLLM deployment, model configuration, benchmarking, batching tuning, GPU operation, and upgrade/rollback, load the vllm tool skill instead. This skill provides the methodology (eval-set design, quantization trade-offs, deployment plans, regression triage); the tool skills own the runbooks.

© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (scripts, references) in ml-engineering of magnus919/agent-skills.

  • SKILL.md
  • README.md
  • evals/evals.json
  • references/evaluation-and-lineage.md
  • references/evaluation.md
  • references/fine-tuning.md
  • references/quantization-inference.md
  • references/training-infrastructure.md
  • scripts/check-eval-overlap.py
  • scripts/test_check_eval_overlap.py
  • templates/drift-response-record.md
  • templates/eval-regression-table.md
  • templates/model-lineage-record.md
  • templates/quantization-decision-record.md
  • templates/training-run-record.md

Open the folder on GitHubat commit 96fbe07

Compare with similar skills

ML Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Engineering compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Engineering this skillmagnus919/agent-skills113—~1.5kAutomated safety check: PassMIT
ML Research LabAnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMIT
Unslothericrisco/rsc-harness167—~3.6kAutomated safety check: PassMIT
Openrlhf TrainingOrchestra-Research/AI-Research-SKILLs13k3 repos~2.1kAutomated safety check: NotesMIT
Quantized Exportwshobson/agents40k—~2kAutomated safety check: PassMIT
Open Weightsericrisco/rsc-harness167—~4.1kAutomated safety check: PassMIT

Similar skills

  • ML Research Lab

    AnastasiyaW/codex-claude-code-config

    Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

    154 GitHub stars~794 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Unsloth

    ericrisco/rsc-harness

    A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

    167 GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Openrlhf Training

    Orchestra-Research/AI-Research-SKILLs

    High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.

    13k GitHub starsUsed in 3 repos~2.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • Quantized Export

    wshobson/agents

    Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

    40k GitHub stars~2k tokensUpdated 4 days ago
    AI & LLM EngineeringAuto-check passed
  • Open Weights

    ericrisco/rsc-harness

    A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

    167 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 3 repos~3k tokens
    AI & LLM EngineeringAuto-check passed

More from magnus919/agent-skills

All 115 skills in this repo
  • Artifact Pyramids

    magnus919/agent-skills

    Organize durable agent research outputs as summaries, analysis, and evidence dossiers.

    113 GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed
  • Ascii City Engine

    magnus919/agent-skills

    Build portable, first-person colored ASCII city engines and small GIS-derived city packs.

    113 GitHub stars~1.1k tokensUpdated 2 days ago
    Auto-check passed
  • Color Management

    magnus919/agent-skills

    Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.

    113 GitHub stars~2.6k tokensUpdated 2 days ago
    Auto-check: notes
  • Data Scientist

    magnus919/agent-skills

    A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…

    113 GitHub stars~4.1k tokensUpdated 2 days ago
    Auto-check passed
  • Docker Compose

    magnus919/agent-skills

    Use Docker Compose to define, run, debug, and harden multi-container applications.

    113 GitHub stars~2k tokensUpdated 2 days ago
    Auto-check: notes
  • Fpga Development

    magnus919/agent-skills

    Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.

    113 GitHub stars~2.7k tokensUpdated 2 days ago
    Auto-check passed

Works with

Questions about ML Engineering

What does ML Engineering do?

Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…. ML Engineering is an agent skill from magnus919/agent-skills. Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity, drift response, and regression triage, grounded in practical engineering patterns for production ML systems.

When should I use ML Engineering?

ML Engineering fits situations like: statistical modeling and experimental design (thats the data scientist); for operating a specific inference engine (thats a tool skill such as llama-cpp.

How do I install ML Engineering in Claude Code?

Run `npx skills add magnus919/agent-skills --skill ml-engineering -a claude-code`. Or copy the skill folder (ml-engineering in magnus919/agent-skills) into .claude/skills/ml-engineering in your project. Claude Code loads it when a task matches its description.

How do I install ML Engineering in Codex?

Run `npx skills add magnus919/agent-skills --skill ml-engineering -a codex`. Or copy the skill folder (ml-engineering in magnus919/agent-skills) into .agents/skills/ml-engineering in your project. Codex loads it when a task matches its description.

Can I use ML Engineering in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill ml-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-engineering, .gemini/skills/ml-engineering, .github/skills/ml-engineering and .opencode/skills/ml-engineering in your project.

What does ML Engineering need to run?

Going by SKILL.md and its folder, ML Engineering needs Python for the scripts in its folder. Our summary lists: Python 3.

Does ML Engineering access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML Engineering safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does ML Engineering use?

ML Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Engineering use?

About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.

What are the alternatives to ML Engineering?

Skills that share tags, products or a category with ML Engineering: ML Research Lab (AnastasiyaW/codex-claude-code-config, 154 stars), Unsloth (ericrisco/rsc-harness, 167 stars), Openrlhf Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Quantized Export (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Engineering?

magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 113 GitHub stars. The repository holds 115 skills in this directory. The repository was last updated on October 6, 2026.

Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.