ML Research Lab
AnastasiyaW/codex-claude-code-config
Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.
Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…
$ npx skills add magnus919/agent-skills --skill ml-engineering -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install magnus919/agent-skills ml-engineering --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/ml-engineering .claude/skills/ml-engineering && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ml-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/ml-engineering into .claude/skills/ml-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-engineering", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/magnus919/agent-skills/tree/main/ml-engineeringType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add magnus919/agent-skills --skill ml-engineering -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install magnus919/agent-skills ml-engineering --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .agents/skills && cp -r skills-src/ml-engineering .agents/skills/ml-engineering && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ml-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/ml-engineering into .agents/skills/ml-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-engineering", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill ml-engineering -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install magnus919/agent-skills ml-engineering --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/ml-engineering .cursor/skills/ml-engineering && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ml-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/ml-engineering into .cursor/skills/ml-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-engineering", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/magnus919/agent-skills.git --path ml-engineering--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add magnus919/agent-skills --skill ml-engineering -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install magnus919/agent-skills ml-engineering --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/ml-engineering .gemini/skills/ml-engineering && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ml-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/ml-engineering into .gemini/skills/ml-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-engineering", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install magnus919/agent-skills ml-engineeringInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add magnus919/agent-skills --skill ml-engineering -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .github/skills && cp -r skills-src/ml-engineering .github/skills/ml-engineering && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ml-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/ml-engineering into .github/skills/ml-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-engineering", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add magnus919/agent-skills --skill ml-engineering -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install magnus919/agent-skills ml-engineering --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/magnus919/agent-skills.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/ml-engineering .opencode/skills/ml-engineering && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ml-engineering" agent skill from https://github.com/magnus919/agent-skills/tree/main/ml-engineering into .opencode/skills/ml-engineering/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ml-engineering", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ml-engineeringPlan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…
ML Engineering is an agent skill from magnus919/agent-skills. Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity, drift response, and regression triage, grounded in practical engineering patterns for production ML systems. Do not use for statistical modeling and experimental design (that's the data scientist) or for operating a specific inference engine (that's a tool skill such as llama-cpp or vllm).
Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 18 other files, including scripts and reference files (for example `README.md`, `evals/evals.json` and `references/evaluation-and-lineage.md`).
It sits in AI & LLM Engineering, covering Fine-tuning and LLM inference and serving. It works with llama.cpp and vLLM. The repository describes itself as: Curated collection of AI agent skills for Hermes and other agent frameworks. The licence is MIT.
Read from SKILL.md and the folder at commit 96fbe07. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
From the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
ML Engineering loads about 1.5k tokens when it runs, and up to ~15k if it reads all its reference files. Until then it costs about 126 tokens; SKILL.md has 570 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from magnus919/agent-skills at commit 96fbe07, republished under its MIT licence (© magnus919). 570 words, ~1,465 tokens.
.claude/skills/ml-engineering/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.Machine learning engineering is the bridge between model research and production systems. This methodology covers the engineering disciplines needed to train, evaluate, deploy, and maintain ML models reliably.
| You own | You don't own |
|---|---|
| Model training — LoRA/QLoRA fine-tuning, full fine-tuning, distributed training | Statistical modeling and experimental design — that's the data scientist |
| Model evaluation — benchmark suites, custom eval sets, regression testing | Causal inference and hypothesis testing — that's the data scientist |
| Quantization — GGUF, GPTQ, AWQ, bitsandbytes | Training data collection and labeling — that's the data/ML ops team |
| Inference serving — vLLM, llama.cpp, TGI, Triton | Business metrics and KPI definition — that's the product manager |
| Evaluation harness — lm-eval-harness, custom pipelines | Data pipeline architecture — that's the data engineer |
| Model deployment — containerization, versioning, A/B testing | Infrastructure provisioning — that's the platform engineer |
| Reference | When to load |
|---|---|
references/fine-tuning.md | Setting up a LoRA/QLoRA/ full fine-tuning run — data prep, hyperparameters, validation strategy |
references/evaluation.md | Evaluating a model — benchmark selection, custom eval sets, regression tracking, comparison methodology |
references/evaluation-and-lineage.md | Metric/configuration decisions, repeated stochastic comparisons, end-to-end lineage, temporal feature parity, drift response, and adaptation/serving tradeoffs |
references/quantization-inference.md | Quantizing a model and serving it — GGUF/GPTQ/AWQ/bitsandbytes comparison, calibration data strategies, KV cache quantization, vLLM/llama.cpp/TGI/Triton architecture, production considerations |
references/training-infrastructure.md | Selecting and provisioning training infrastructure — GPU selection, VRAM budgeting, multi-GPU strategies (DDP/FSDP/DeepSpeed), cloud vs on-prem, storage, monitoring |
| Template | When to Use |
|---|---|
templates/training-run-record.md | Recording a training or fine-tuning run — model and data versions, full config, environment, eval results — so it can be reproduced |
templates/eval-regression-table.md | Tracking model quality across runs and triaging a regression — one row per eval case or capability subset |
templates/quantization-decision-record.md | Recording a quantization decision — baseline, candidates compared, quality threshold, and rollback path |
templates/model-lineage-record.md | Linking data/features, code/configuration, runs, artifacts, evaluations, registry state, and deployed serving versions |
templates/drift-response-record.md | Recording drift signals, thresholds, diagnosis, retrain/rollback decisions, and post-action evidence |
| Script | When to Use |
|---|---|
scripts/check-eval-overlap.py | Checking a training corpus against an eval corpus for test-set leakage (shared n-grams); --json for CI, exit 1 when an eval file exceeds the overlap threshold |
evals/evals.json — output-quality eval manifest for this skill: fine-tuning plan review, eval-set design, quantization decision, deployment plan, regression triage, and training-run reproducibility.
Measure before you optimize — Never quantize, prune, or distill a model without first measuring its baseline performance. Optimization without measurement is guessing.
Reproducibility is non-negotiable — Every training run needs a reproducible config: seed, data version, hyperparameters, and evaluation methodology. If you can't reproduce it, you can't ship it.
Baseline first — Before running an expensive fine-tuning run, establish a baseline with the base model. If the base model is already good enough, the fine-tuning budget is better spent elsewhere.
Test at the boundary — Model evaluation is most informative at the edges of the capability distribution, not at the center. Hard examples reveal more than easy ones.
The evaluation set is a liability — Every example in your eval set is a potential test-set leak. Use held-out sets, rotate examples, and periodically audit for contamination with the overlap checker.
Do not use this skill for statistical modeling, experimental design, or causal inference — that's the data scientist's discipline. Do not use it to operate a specific inference engine: for llama.cpp installation, model loading, benchmarking, and troubleshooting, load the llama-cpp tool skill instead; for vLLM deployment, model configuration, benchmarking, batching tuning, GPU operation, and upgrade/rollback, load the vllm tool skill instead. This skill provides the methodology (eval-set design, quantization trade-offs, deployment plans, regression triage); the tool skills own the runbooks.
© magnus919, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 14 other files (scripts, references) in ml-engineering of magnus919/agent-skills.
Open the folder on GitHubat commit 96fbe07
ML Engineering next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| ML Engineering this skillmagnus919/agent-skills | 113 | — | ~1.5k | Automated safety check: Pass | MIT | |
| ML Research LabAnastasiyaW/codex-claude-code-config | 154 | — | ~794 | Automated safety check: Pass | MIT | |
| Unslothericrisco/rsc-harness | 167 | — | ~3.6k | Automated safety check: Pass | MIT | |
| Openrlhf TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.1k | Automated safety check: Notes | MIT | |
| Quantized Exportwshobson/agents | 40k | — | ~2k | Automated safety check: Pass | MIT | |
| Open Weightsericrisco/rsc-harness | 167 | — | ~4.1k | Automated safety check: Pass | MIT |
AnastasiyaW/codex-claude-code-config
Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.
ericrisco/rsc-harness
A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…
Orchestra-Research/AI-Research-SKILLs
High-performance RLHF framework with Ray+vLLM acceleration. An agent skill from Orchestra-Research/AI-Research-SKILLs.
wshobson/agents
Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.
ericrisco/rsc-harness
A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
magnus919/agent-skills
Organize durable agent research outputs as summaries, analysis, and evidence dossiers.
magnus919/agent-skills
Build portable, first-person colored ASCII city engines and small GIS-derived city packs.
magnus919/agent-skills
Manage color workflows with ICC profiles, working spaces, gamut mapping, and color science.
magnus919/agent-skills
A skill your agent uses for PhD-level expertise in data science, statistics, and machine learning: rigorous statistical analysis, experimental design, causal inference, advanced modeling, research…
magnus919/agent-skills
Use Docker Compose to define, run, debug, and harden multi-container applications.
magnus919/agent-skills
Design, review, simulate, and verify FPGA logic using explicit RTL contracts, clock and reset models, CDC analysis, timing constraints, and reproducible implementation evidence.
Categories
Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…. ML Engineering is an agent skill from magnus919/agent-skills. Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity, drift response, and regression triage, grounded in practical engineering patterns for production ML systems.
ML Engineering fits situations like: statistical modeling and experimental design (thats the data scientist); for operating a specific inference engine (thats a tool skill such as llama-cpp.
Run `npx skills add magnus919/agent-skills --skill ml-engineering -a claude-code`. Or copy the skill folder (ml-engineering in magnus919/agent-skills) into .claude/skills/ml-engineering in your project. Claude Code loads it when a task matches its description.
Run `npx skills add magnus919/agent-skills --skill ml-engineering -a codex`. Or copy the skill folder (ml-engineering in magnus919/agent-skills) into .agents/skills/ml-engineering in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add magnus919/agent-skills --skill ml-engineering -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-engineering, .gemini/skills/ml-engineering, .github/skills/ml-engineering and .opencode/skills/ml-engineering in your project.
Going by SKILL.md and its folder, ML Engineering needs Python for the scripts in its folder. Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
ML Engineering is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.5k tokens (SKILL.md is roughly 5.9k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 13k tokens, read only when the agent opens those files.
Skills that share tags, products or a category with ML Engineering: ML Research Lab (AnastasiyaW/codex-claude-code-config, 154 stars), Unsloth (ericrisco/rsc-harness, 167 stars), Openrlhf Training (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Quantized Export (wshobson/agents, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
magnus919 (a GitHub user) maintains it in magnus919/agent-skills, which has 113 GitHub stars. The repository holds 115 skills in this directory. The repository was last updated on October 6, 2026.
Source: magnus919/agent-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.