Run the CLOSED training loop on a MOSS/duck policy without a human having to ask "is it done yet": launch through the lab, block until the run finishes, print the standard report and the per-term…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Train Loop

skills CLI
$ npx skills add jonathanhawkins/microduck-lab --skill train-loop -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jonathanhawkins/microduck-lab train-loop --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jonathanhawkins/microduck-lab.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/train-loop .claude/skills/train-loop && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
train-loop
GitHub stars
129
Token cost
~947 tokens
SKILL.md length
453 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Run the CLOSED training loop on a MOSS/duck policy without a human having to ask "is it done yet": launch through the lab, block until the run finishes, print the standard report and the per-term…

  • A training run is started and its result matters
  • SKILL.md covers The three rules this loop… and Adjusting, then relaunching
  • Calls uv and curl
  • : train it again

What it does

Train Loop is an agent skill from jonathanhawkins/microduck-lab. Run the CLOSED training loop on a MOSS/duck policy without a human having to ask "is it done yet": launch through the lab, block until the run finishes, print the standard report and the per-term reward budget, then adjust and relaunch. Use whenever a training run is started and its result matters. Trigger on: "train it again", "when it's done check the results", "tune the rules and retrain", "let me know when training finishes", "is it still training".

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. The repository describes itself as: Train RL policies for the Pollen Microduck 🦆 on an ordinary Mac, no CUDA GPU, and watch them learn live in the browser. The licence is Apache-2.0.

When your agent uses it

  • A training run is started and its result matters
  • : train it again
  • Its done check the results
  • Tune the rules and retrain

Example prompts

  • “is it done yet”
  • “train it again”
  • “when it”
  • “/train-loop”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit bbf0326. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • curl

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Train Loop loads about 947 tokens when it runs. Until then it costs about 117 tokens; SKILL.md has 453 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~117
When it runs · the whole SKILL.md, loaded when a task matches
~947

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jonathanhawkins/microduck-lab at commit bbf0326, republished under its Apache-2.0 licence (© jonathanhawkins). 453 words, ~947 tokens.

Download SKILL.mdSave it as .claude/skills/train-loop/SKILL.md (or your agent's skills folder).
name
train-loop
description
Run the CLOSED training loop on a MOSS/duck policy without a human having to ask "is it done yet": launch through the lab, block until the run finishes, print the standard report and the per-term reward budget, then adjust and relaunch. Use whenever a training run is started and its result matters. Trigger on: "train it again", "when it's done check the results", "tune the rules and retrain", "let me know when training finishes", "is it still training".

train-loop — launch, wait, measure, adjust, relaunch

A human should not have to poll an agent for "is it done". Run the waiter in the BACKGROUND: it exits when the run ends, which re-invokes the agent with the report already computed.

bash
uv run python scripts/watch_training.py 127.0.0.1:8788 40

It blocks on GET /teach/status, then exports the finished run (never judge a raw checkpoint — export-walk bakes the normalizer in) and prints:

  • scripts/pick_report.py — picks, can displacement, base command, chassis turn, floor dragging, torque saturation, and the wrist-vs-axis SLOPE with its p-value.
  • scripts/reward_budget.py — what each reward term is worth per episode.

The three rules this loop exists to enforce

1. Evaluate a policy under the flags it TRAINED with. eval_env_kwargs(run) reads them off run.json — what the policy SAW (OBS_FLAGS) and what it was ALLOWED TO DO (DYN_FLAGS). Getting this wrong does not look like an error, it looks like a RESULT:

the mistakewhat it reportedthe truth
attitude slots filled for an axis-blind policy0/12 picks10/12
publish_size omitted from the guard3 grips63
base_lock omitted from the guard23.7 deg of chassis turn0.3 deg

All three happened on 2026-09-25, and the last two happened AFTER a guard was written — because the guard covered some flags and read as covering all of them. A new flag that changes the observation goes in OBS_FLAGS; one that changes the dynamics goes in DYN_FLAGS.

2. Read the reward budget after ANY change to what a term measures. Retargeting the progress term from the chassis to the gripper without rescaling dropped it from ~+18 per episode to +1.47, and two full runs trained with almost no approach shaping before anyone noticed. Changing what a term MEASURES changes its MAGNITUDE.

3. Check a new knob's reachable set before believing in it. Print the share of ticks it fires on. A rung calibrated under different semantics from where it was applied fired on 1 of 40 episodes instead of the predicted 15%.

Show full SKILL.md (135 more words)Show less

Adjusting, then relaunching

Knobs reach the env as MICRODUCK_MOSS_* in the /teach request's env, and Body.train_env_kwargs copies them into run.json so the run records what it trained on. Add every new knob there, or the run is unreproducible and the loop above cannot evaluate it correctly.

bash
curl -s -X POST http://127.0.0.1:8788/teach -H 'Content-Type: application/json' \
  -d '{"text":"pick up the can","robot":"moss","initFrom":"<donor>",
       "env":{"MICRODUCK_MOSS_PICK_RUNG":"2","MICRODUCK_MOSS_BASE_LOCK":"1"},
       "steps":800000,"weights":{}}'

Restart the lab BEFORE launching if any env module changed (the stage previews in-process; the trainer is a subprocess), and confirm the trainee is really on stage — see the two sections in microduck_local/AGENTS.md.

Warm starting into a NEW observation slot needs the normalizer reseeded first: a donor trained with a slot dead carries variance ~1e-10 there, so a real value arrives as a z-score in the thousands and the policy stops working. Copy the donor run, overwrite obs_rms.mean/var for those slots with the MEASURED distribution, warm-start from the copy.

© jonathanhawkins, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/train-loop of jonathanhawkins/microduck-lab.

Open the folder on GitHubat commit bbf0326

Compare with similar skills

Train Loop next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Train Loop compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Train Loop this skilljonathanhawkins/microduck-lab129—~947Automated safety check: PassApache-2.0
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
Add Oponnx/onnx22k—~1.2kAutomated safety check: PassApache-2.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Add Function Bodyonnx/onnx22k—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Op

    onnx/onnx

    Add a new ONNX operator or update an existing operator to a new opset version.

    22k GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add a function body definition to an ONNX operator, defining how it decomposes into simpler ops.

    22k GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Distributed

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

    24k GitHub stars~660 tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from jonathanhawkins/microduck-lab

  • Record World

    jonathanhawkins/microduck-lab

    Record a /sim WORLD scenario — the living room, the playroom tidy loop, a soccer pitch, any scenario JSON — to an mp4, a captioned contact sheet and an events log, headless and under a seed, then…

    129 GitHub stars~1.2k tokensUpdated 6 days ago
    Auto-check passed
  • Render Rollout

    jonathanhawkins/microduck-lab

    Render a microduck policy rollout to video AND to a frame contact sheet with per-frame diagnostics burned in, then READ the sheet to see what the policy actually does.

    129 GitHub stars~3.2k tokensUpdated 6 days ago
    Auto-check passed
  • Tidy Trace

    jonathanhawkins/microduck-lab

    Debug the playroom tidy loop (Track 12) — trace one run state by state, see every release, landing and fall with context, and re-measure the walker facts the brain's constants rest on.

    129 GitHub stars~1.1k tokensUpdated 6 days ago
    Auto-check passed
  • Restart Servers

    jonathanhawkins/microduck-lab

    Restart the microduck dev stack — the duck-lab backend (farm, :8788) and the duck-viewer dev server (:63317).

    129 GitHub stars~1.1k tokensUpdated 6 days ago
    Auto-check passed
  • Sim Smoke

    jonathanhawkins/microduck-lab

    Look at the /sim world page the way a user would — bring up the lab in world mode and the viewer, open the page in headless Chromium, press keys, screenshot it, and READ the screenshot and console.

    129 GitHub stars~1.2k tokensUpdated 6 days ago
    Auto-check passed
  • Watch Training

    jonathanhawkins/microduck-lab

    Look at what the ACTIVE teach run is practicing right now: renders the live checkpoint under the trainer's own env knobs (actuator, spawn mix, reward gates — read from the live trainer process) and…

    129 GitHub stars~525 tokensUpdated 6 days ago
    Auto-check passed

Questions about Train Loop

What does Train Loop do?

Run the CLOSED training loop on a MOSS/duck policy without a human having to ask "is it done yet": launch through the lab, block until the run finishes, print the standard report and the per-term…. Train Loop is an agent skill from jonathanhawkins/microduck-lab. Run the CLOSED training loop on a MOSS/duck policy without a human having to ask "is it done yet": launch through the lab, block until the run finishes, print the standard report and the per-term reward budget, then adjust and relaunch.

When should I use Train Loop?

Train Loop fits situations like: A training run is started and its result matters; : train it again; its done check the results; tune the rules and retrain.

How do I install Train Loop in Claude Code?

Run `npx skills add jonathanhawkins/microduck-lab --skill train-loop -a claude-code`. Or copy the skill folder (.claude/skills/train-loop in jonathanhawkins/microduck-lab) into .claude/skills/train-loop in your project. Claude Code loads it when a task matches its description.

How do I install Train Loop in Codex?

Run `npx skills add jonathanhawkins/microduck-lab --skill train-loop -a codex`. Or copy the skill folder (.claude/skills/train-loop in jonathanhawkins/microduck-lab) into .agents/skills/train-loop in your project. Codex loads it when a task matches its description.

Can I use Train Loop in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jonathanhawkins/microduck-lab --skill train-loop -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/train-loop, .gemini/skills/train-loop, .github/skills/train-loop and .opencode/skills/train-loop in your project.

What does Train Loop need to run?

Going by SKILL.md and its folder, Train Loop needs the command-line tools its instructions call (uv and curl). Our summary lists: Python 3.

Does Train Loop access the network?

SKILL.md contains no URLs. Its commands use uv and curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Train Loop safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Train Loop use?

Train Loop is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Train Loop use?

About 947 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Train Loop?

Skills that share tags, products or a category with Train Loop: Add Uint Support (pytorch/pytorch, 104k stars), Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Op (onnx/onnx, 22k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Train Loop?

jonathanhawkins (a GitHub user) maintains it in jonathanhawkins/microduck-lab, which has 129 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 1, 2026.

Source: jonathanhawkins/microduck-lab on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.