Agent skill

ML Research Lab

by AnastasiyaW in AnastasiyaW/codex-claude-code-config

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

MITAuto-check passedAI & LLM Engineering

Install ML Research Lab

skills CLI
$ npx skills add AnastasiyaW/codex-claude-code-config --skill ml-research-lab -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AnastasiyaW/codex-claude-code-config ml-research-lab --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AnastasiyaW/codex-claude-code-config.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ai-ml/ml-research-lab .claude/skills/ml-research-lab && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ml-research-lab
GitHub stars
154
Token cost
~794 tokens
SKILL.md length
352 words
Files
1
Skills in repo
50
Repo updated
First seen
Licence
MIT

At a glance

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

  • Works in 7 steps: Freeze the question as a measurable… → Identify dataset provenance, labels,… → Pick the smallest baseline that can… → …
  • Working on ML experiments
  • SKILL.md covers Operating Loop, Domain Routing, Verification Gates and Adoption Boundary
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

ML Research Lab is an agent skill from AnastasiyaW/codex-claude-code-config. Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability. Use when working on ML experiments, training data, model benchmarks, RunPod/GPU runs, classifier quality, vLLM/GGUF serving, SHAP-style model explanations, or research-to-code iterations. Do not use for a simple code edit that has no ML dataset, metric, model, or experiment artifact.

Its SKILL.md is about 790 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Machine learning, Fine-tuning and LLM inference and serving. It works with llama.cpp and vLLM. The repository describes itself as: Claude Code, Codex, and multi-agent configuration system: principles, hooks, skills, and workflow patterns for AI-assisted development. The licence is MIT.

When your agent uses it

  • Working on ML experiments
  • Model benchmarks
  • RunPod/GPU runs
  • Classifier quality

Example prompts

  • “/ml-research-lab”

Workflow steps

7 steps, taken from the first numbered list in SKILL.md.

  1. Freeze the question as a measurable hypothesis.
  2. Identify dataset provenance, labels, splits, leakage risks, and regeneration cost.
  3. Pick the smallest baseline that can disprove the idea.
  4. Define metrics before training. For release claims, require train/val/test split,
  5. Run or wire experiment tracking before long jobs start.
  6. Save artifacts: config, command, data manifest, metrics JSON/CSV, logs, model hash,
  7. Compare against baseline, then keep/discard the change from evidence.

What it can do on your machine

Read from SKILL.md and the folder at commit 67709af. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

ML Research Lab loads about 794 tokens when it runs. Until then it costs about 112 tokens; SKILL.md has 352 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~112
When it runs · the whole SKILL.md, loaded when a task matches
~794

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AnastasiyaW/codex-claude-code-config at commit 67709af, republished under its MIT licence (© AnastasiyaW). 352 words, ~794 tokens.

Download SKILL.mdSave it as .claude/skills/ml-research-lab/SKILL.md (or your agent's skills folder).
name
ml-research-lab
description
Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability. Use when working on ML experiments, training data, model benchmarks, RunPod/GPU runs, classifier quality, vLLM/GGUF serving, SHAP-style model explanations, or research-to-code iterations. Do not use for a simple code edit that has no ML dataset, metric, model, or experiment artifact.

ML Research Lab

Use this skill as the compact router for ML work. It is derived from an audit of synthetic-sciences/openscience at commit 531467c, but does not require running OpenScience or loading its full 250+ skill set.

Operating Loop

  1. Freeze the question as a measurable hypothesis.
  2. Identify dataset provenance, labels, splits, leakage risks, and regeneration cost.
  3. Pick the smallest baseline that can disprove the idea.
  4. Define metrics before training. For release claims, require train/val/test split, no test-set model selection, and multi-seed proof when cost permits.
  5. Run or wire experiment tracking before long jobs start.
  6. Save artifacts: config, command, data manifest, metrics JSON/CSV, logs, model hash, and a short conclusion.
  7. Compare against baseline, then keep/discard the change from evidence.

Domain Routing

  • Dataset or scrape cleanup: start from data quality, deduplication, leakage checks, train/eval splits, and regeneration notes.
  • Classical classifier or tabular baseline: use scikit-learn-style pipelines with preprocessing inside the pipeline and stratified splits for classification.
  • Model debugging or trust: add SHAP/explainability for feature importance, leakage, bias/proxy features, and misclassified samples.
  • LLM fine-tuning: prefer JSONL chat format, data validation, LoRA/QLoRA baseline, and tracked runs before scaling.
  • Single-GPU fast LoRA/QLoRA: consider Unsloth only after checking hardware, CUDA, model support, and export target.
  • Large or production inference: use vLLM for high-throughput GPU serving, GGUF or llama.cpp for local/Apple/CPU-friendly deployment, and TensorRT-LLM only when the NVIDIA production optimization cost is justified.
  • Research write-up: report method, dataset, exact metric formula, baseline source, limitations, and failure cases.
Show full SKILL.md (105 more words)Show less

Verification Gates

  • Data gate: schema valid, duplicates/leakage checked, split manifest saved.
  • Metric gate: exact metric formula named; if benchmarked, original baseline source and benchmark code checked.
  • Runtime gate: command/log path and environment captured; GPU memory and errors checked for long runs.
  • Tracking gate: metrics are retrievable as JSON/CSV or a dashboard link plus local export.
  • Deployment gate: latency, throughput, memory, and OOM behavior measured before claiming production readiness.

Adoption Boundary

Do not import broad external skill collections wholesale. Use the inventory script scripts/openscience_skill_inventory.py to rank candidates, inspect the relevant source skill manually, then promote only compact workflows or deterministic scripts that improve our own tests.

© AnastasiyaW, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ai-ml/ml-research-lab of AnastasiyaW/codex-claude-code-config.

Open the folder on GitHubat commit 67709af

Compare with similar skills

ML Research Lab next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

ML Research Lab compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
ML Research Lab this skillAnastasiyaW/codex-claude-code-config154—~794Automated safety check: PassMIT
Unslothericrisco/rsc-harness180—~3.6kAutomated safety check: PassMIT
ML Engineeringmagnus919/agent-skills115—~1.5kAutomated safety check: PassMIT
Quantized Exportwshobson/agents40k—~2kAutomated safety check: PassMIT
Open Weightsericrisco/rsc-harness180—~4.1kAutomated safety check: PassMIT
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT

Similar skills

  • Unsloth

    ericrisco/rsc-harness

    A skill your agent uses when fine-tuning an open-weight LLM fast on ONE GPU with low VRAM — Unsloth's fast model loaders with 4-bit QLoRA and the trl trainer, response-only loss masking so the…

    180 GitHub stars~3.6k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • ML Engineering

    magnus919/agent-skills

    Plan and execute production ML engineering work — model training and fine-tuning (LoRA/QLoRA), evaluation and eval-set design, quantization decisions, inference deployment, lineage, feature parity…

    115 GitHub stars~1.5k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Quantized Export

    wshobson/agents

    Export a promoted fine-tuned model in the right deployment format — merged safetensors, LoRA-only, GGUF with imatrix, or FP8.

    40k GitHub stars~2k tokensUpdated 6 days ago
    AI & LLM EngineeringAuto-check passed
  • Open Weights

    ericrisco/rsc-harness

    A skill your agent uses when choosing an open-weight LLM and clearing it for use — which family and size fit the task, the hardware and the budget, and above all whether the license permits shipping.

    180 GitHub stars~4.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Quantization

    vllm-project/vllm-omni

    Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from AnastasiyaW/codex-claude-code-config

All 50 skills in this repo
  • Bug Reproducer

    AnastasiyaW/codex-claude-code-config

    Find likely software bugs in a codebase, rank concrete bug candidates, and prove or reject them with focused regression tests before proposing a fix.

    154 GitHub stars~4.1k tokensUpdated today
    Auto-check passed
  • Motion Framer

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when implementing Motion or Framer Motion in React/JavaScript: interactive UI components, micro-interactions, gestures, layout or page transitions, and scroll-based animation.

    154 GitHub starsUsed in 1 repo~5.2k tokens
    Auto-check passed
  • Proof Verify

    AnastasiyaW/codex-claude-code-config

    Plan-based verification - freeze acceptance criteria before building, then verify after with an independent fresh-context agent (the builder must not verify their own work).

    154 GitHub stars~2.6k tokensUpdated today
    Auto-check passed
  • Workflow Orchestration

    AnastasiyaW/codex-claude-code-config

    Написание и запуск Claude Code dynamic workflows (JS-оркестратор субагентов).

    154 GitHub stars~3.8k tokensUpdated today
    Auto-check passed
  • Notebooklm Grounded Research

    AnastasiyaW/codex-claude-code-config

    A skill your agent uses when: NotebookLM, notebooklm MCP, large documentation sets, courses, books, papers, or citation-backed research are mentioned.

    154 GitHub stars~2.4k tokensUpdated today
    Auto-check: warnings
  • Deepseek Provider Contract

    AnastasiyaW/codex-claude-code-config

    Validate a proposed DeepSeek API integration before any key or project context is sent: check thinking-mode tool-call history, strict-schema assumptions, bounded output, and provider data boundaries.

    154 GitHub stars~1.2k tokensUpdated today
    Auto-check passed

Works with

Questions about ML Research Lab

What does ML Research Lab do?

Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability. ML Research Lab is an agent skill from AnastasiyaW/codex-claude-code-config. Machine-learning research loop for dataset curation, fine-tuning, evaluation, inference deployment, experiment tracking, and model explainability.

When should I use ML Research Lab?

ML Research Lab fits situations like: working on ML experiments; model benchmarks; runPod/GPU runs; classifier quality.

How do I install ML Research Lab in Claude Code?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill ml-research-lab -a claude-code`. Or copy the skill folder (skills/ai-ml/ml-research-lab in AnastasiyaW/codex-claude-code-config) into .claude/skills/ml-research-lab in your project. Claude Code loads it when a task matches its description.

How do I install ML Research Lab in Codex?

Run `npx skills add AnastasiyaW/codex-claude-code-config --skill ml-research-lab -a codex`. Or copy the skill folder (skills/ai-ml/ml-research-lab in AnastasiyaW/codex-claude-code-config) into .agents/skills/ml-research-lab in your project. Codex loads it when a task matches its description.

Can I use ML Research Lab in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AnastasiyaW/codex-claude-code-config --skill ml-research-lab -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ml-research-lab, .gemini/skills/ml-research-lab, .github/skills/ml-research-lab and .opencode/skills/ml-research-lab in your project.

What does ML Research Lab need to run?

SKILL.md names no scripts, command-line tools or credentials: ML Research Lab is instructions for the agent only.

Does ML Research Lab access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is ML Research Lab safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does ML Research Lab use?

ML Research Lab is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does ML Research Lab use?

About 794 tokens (SKILL.md is roughly 3.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to ML Research Lab?

Skills that share tags, products or a category with ML Research Lab: Unsloth (ericrisco/rsc-harness, 180 stars), ML Engineering (magnus919/agent-skills, 115 stars), Quantized Export (wshobson/agents, 40k stars) and Open Weights (ericrisco/rsc-harness, 180 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains ML Research Lab?

AnastasiyaW (a GitHub user) maintains it in AnastasiyaW/codex-claude-code-config, which has 154 GitHub stars. The repository holds 50 skills in this directory. The repository was last updated on October 9, 2026.

Source: AnastasiyaW/codex-claude-code-config on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.