Shipping and Launch Checklist
addyosmani/agent-skills
Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.
Preflight and launch a Lego-RL run (train, eval or infer). An agent skill from LegoX/Lego-RL.
$ npx skills add LegoX/Lego-RL --skill run -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install LegoX/Lego-RL run --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/run .claude/skills/run && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "run" agent skill from https://github.com/LegoX/Lego-RL/tree/main/.claude/plugins/rl-plugin/skills/run into .claude/skills/run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/LegoX/Lego-RL/tree/main/.claude/plugins/rl-plugin/skills/runType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add LegoX/Lego-RL --skill run -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install LegoX/Lego-RL run --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/run .agents/skills/run && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "run" agent skill from https://github.com/LegoX/Lego-RL/tree/main/.claude/plugins/rl-plugin/skills/run into .agents/skills/run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LegoX/Lego-RL --skill run -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install LegoX/Lego-RL run --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/run .cursor/skills/run && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "run" agent skill from https://github.com/LegoX/Lego-RL/tree/main/.claude/plugins/rl-plugin/skills/run into .cursor/skills/run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/LegoX/Lego-RL.git --path .claude/plugins/rl-plugin/skills/run--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add LegoX/Lego-RL --skill run -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install LegoX/Lego-RL run --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/run .gemini/skills/run && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "run" agent skill from https://github.com/LegoX/Lego-RL/tree/main/.claude/plugins/rl-plugin/skills/run into .gemini/skills/run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install LegoX/Lego-RL runInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add LegoX/Lego-RL --skill run -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/run .github/skills/run && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "run" agent skill from https://github.com/LegoX/Lego-RL/tree/main/.claude/plugins/rl-plugin/skills/run into .github/skills/run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add LegoX/Lego-RL --skill run -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install LegoX/Lego-RL run --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/run .opencode/skills/run && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "run" agent skill from https://github.com/LegoX/Lego-RL/tree/main/.claude/plugins/rl-plugin/skills/run into .opencode/skills/run/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "run", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
runPreflight and launch a Lego-RL run (train, eval or infer). An agent skill from LegoX/Lego-RL.
Run is an agent skill from LegoX/Lego-RL. Preflight and launch a Lego-RL run (train, eval or infer). Runs /rl:check internally and refuses on any blocking failure, shows the fully resolved run parameters and waits for explicit confirmation, then launches the runner in the background by default — these runs take hours, so a foreground default would hold the session hostage. For multi-node train/infer it prints the exact per-node command instead of SSH-ing anywhere. Before launching it states where the run's log will actually land and whether the running…
Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Lego-RL: Harness-Native Reinforcement Learning for Coding Agents. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 7c30234. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
bashcurlkindFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Run loads about 3k tokens when it runs. Until then it costs about 202 tokens; SKILL.md has 1,166 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from LegoX/Lego-RL at commit 7c30234, republished under its Apache-2.0 licence (© LegoX). 1,166 words, ~3,013 tokens.
.claude/skills/run/SKILL.md (or your agent's skills folder).Preflight, review the parameters, confirm, then launch in the background. Refuses if any check blocks, or if a run is already live on this box.
Same as /rl:check Step 0: find the repo root, resolve the config
(path / bare name / ask when absent or ambiguous), infer the kind
(train | eval | infer). Never guess which config the user meant — a
wrong guess here burns a cluster for hours.
bash scripts/lib/live_probe.sh <kind> 2>&1 | grep -E '^WARN +job:'If any job:trainer or job:runner line is alive, abort:
A run is already in flight; refusing to start another:
pids: <pid + etime, from ps -p <pid> -o pid=,comm=,etime=>
log: <most recent *.log under logs/>
Let it finish, or stop it yourself (kill -INT <pid>), then re-run /rl:run.A bare job:vllm with no runner above it is a foreign serving job, not
ours — that is Step 2's business (it blocks on GPU ownership), not an abort
message about "our" run.
Run the check skill's logic on this config (invoke that skill, or inline its
Steps 1–3). If the verdict is SAFE TO RUN: ❌ NO, abort with the same
consolidated report, prefixed:
Preflight failed — fix the items below before /rl:run.Do not invent values, skip a check, or pass a force flag; there is no force
flag. If the user says "just launch it anyway", explain exactly which check blocks
and what it costs to ignore it (each ✗ FATAL maps to a real past incident, written
up in the Troubleshooting section of the docs site), then ask them to fix the config.
Configs are user-owned; this skill is a launcher, not an editor.
PREFLIGHT_ONLY=1 already printed the run configuration (<kind>) block in
Step 2. Show that block — verbatim — and then draw attention to the handful of
fields that decide whether the run is worth its hours. Keep this short; the
block above already has the detail.
Worth double-checking before this run:
• data <train/val index, or dataset, or shard> ← is this the one you meant?
• scale <N nodes × 8 gpu>, <trials/step>, epochs=<N> ← roughly <estimate> hours/step
• resume <RESUME_FROM, or "auto: picks up the newest ckpt in the exp dir">
• naming project=<...> exp=<...>
• logs <the real TRAIN_LOG path>
dashboard <visible / not visible: it is serving <log-dir>>
• first step <R3 on → pearson ≈ 0.999; otherwise routing misalignment silently
destroys the gradients>Where the log lands, and whether the board will see it — this line is not
cosmetic. scripts/templates/verl/common.env derives
TRAIN_LOG=${HARBOR_LOG_DIR}/${TRAINER_EXPERIMENT_NAME}.log, and a real config
overrides HARBOR_LOG_DIR to a per-experiment directory under the shared trials
root. Only the template default is <repo>/logs. The dashboard globs its
--log-dir one level, no recursion (webui/server.py) — so a config that
overrides HARBOR_LOG_DIR produces a run that trains fine and is invisible on
the board for its whole life unless something links it in. Resolve both sides
before asking for the yes:
bash scripts/train/train.sh --dry-run <config> 2>&1 | grep -E 'train log|vLLM log'
pgrep -af 'server\.py.*--log-dir' # which dirs a board actually servesIf dirname(TRAIN_LOG) is not one of the served dirs — and "no board running at
all" counts as not served — say so and ask. Do not launch silently, and do not
pick for them:
This config writes its training log to
<TRAIN_LOG>
but the dashboard (pid=<P>, port=<P>) is serving
<served log-dir>
One-level glob, no match → this run will never appear on the board
(training itself is unaffected).
[1] symlink it into <served-dir>/<name>.log right after launch
(no dashboard restart — recommended)
[2] leave it, watch wandb only
[3] stop and change HARBOR_LOG_DIR in the config first
Which one?Option 1 is one ln -s in Step 6, after the real exp_name is known. It is the
only choice needing neither a restart nor a config edit, and the symlink name
becomes the run's id in the UI — the analysis panels get their exp-dir mapping
from the log contents, so any readable name works. Never move the log (the
runner holds it open through tee), and never leave a dangling symlink under a
log dir — that breaks /api/runs for every run.
Then ask, in one line:
Launch with this configuration? [Y]/[N] (N = edit the config first)Never launch without an explicit yes. If the user says no, stop cleanly and
tell them to edit the config, then re-run /rl:run.
exp_name gotcha — unless the config pins EXP_NAME, train.sh derives it
as harbor-<scaffold>-<EXP_TAG>-<timestamp> at launch time, so the name in
the preflight output is not the name the run will have. Never report the
preflight-time name as final; read the real one back in Step 6. If the user
needs a stable name (resuming, comparing runs), suggest pinning EXP_NAME= in
the config before launching.
Default = background. These runs last hours; foreground holds the session.
Launch mode? [B]ackground (default: nohup setsid, logs to logs/launch_<ts>.log)
[F]oreground (holds this session until the run ends)Accept B / F / <enter> (= background).
Read topology (train) or vllm topo (infer) out of the parameter block.
NNODES > 1 — ray_bringup.sh branches on local IP: the node whose
IP equals MASTER_ADDR (default: the launching node's first IP) becomes the ray
head and drives training; every other node joins and exits. So the same
command must be run on each of the NNODES nodes.VLLM_NNODES > 1 — same shape, keyed on VLLM_HEAD_HOST; rank-0
serves the API, the others run --headless DP shards.Print the command block and tell the user to run it on the other nodes themselves. Never SSH to another node, and never assume the workers are up just because the head started.
Multi-node: run the command below once on every node (this machine is already covered
by the launch just performed)
nodes <ip2>, <ip3>, ...:
cd <repo_root> && MASTER_ADDR=<head_ip> nohup setsid bash scripts/train/train.sh <config> > logs/launch_<ts>_$(hostname -s).log 2>&1 &
The head node waits for all <NNODES> nodes to join before training starts; a worker that
never came up shows as the head spinning in its ray status poll.Timestamp: TS=$(date -u +%Y%m%d-%H%M%S). Log: logs/launch_${TS}.log
(mkdir -p logs first).
cd <repo_root> && nohup setsid bash scripts/<kind>/<kind>.sh <config> > "logs/launch_${TS}.log" 2>&1 < /dev/null &
PARENT_PID=$!
disownsetsid matters: it detaches the run from this session so a disconnect does not
take the training down with it.
Then, without polling in a tight loop:
Wait up to 60s (say, 5s sleeps) for the launch log to grow past 0 bytes.
Read its first ~120 lines and pull out:
exp= value from the run configuration block,[FATAL].Once they appear, capture the PIDs:
pgrep -f 'scripts/<kind>/<kind>.sh' and, for train,
pgrep -f 'fully_async_main|main_ppo' (the trainer takes a few minutes to
show up — report it as pending rather than waiting for it).
If Step 3 chose [1], link the log into the served dir now that exp_name is
real. Verify with the API rather than assuming — the run only shows up once it
has training steps, so an empty steps right after launch is expected:
ln -sfn "<TRAIN_LOG>" "<served-dir>/<readable-name>.log"
curl -s -m 300 "http://127.0.0.1:<P>/api/runs" | grep -c '<readable-name>'For a multi-node run, link only the ray-head node's log — that is the one
carrying the trainer output; the worker exp dirs hold just
*_train_gpu_wandb.log, and the analysis panels recover the sibling exp dirs
on their own.
If the log stays empty for 60s, or the runner exits non-zero within 60s, print the tail of the launch log and stop. Do not retry automatically and do not "fix" the config to get past it.
cd <repo_root> && bash scripts/<kind>/<kind>.sh <config> 2>&1 | tee "logs/launch_${TS}.log"Stream it; on exit report the exit code and the log tail.
Tight summary, ≤ 15 lines:
Launched (background)
kind / config: <kind> · <config path relative to the repo>
exp_name: <the real name read back from the launch log; if not there yet,
"starting — check the log in 30s">
launch log: logs/launch_<TS>.log
train log: <absolute TRAIN_LOG path, read from the launch log's "train log:" line>
vllm log: <absolute VLLM_LOG path, same source>
trials: <HARBOR_TRIALS_DIR>
dashboard: <URL, run name = <readable-name> / not visible (user chose [2])>
pids: runner=<pid> trainer=<pid or "starting">
stage: vLLM loading weights + cuda-graph capture; ~10–20 min before step 1
Watch:
tail -F logs/launch_<TS>.log
nvidia-smi
/rl:status # once it is running, check that the first step is healthy
Stop:
kill -INT <runner_pid> # let it wind down; do not kill -9For train, close with the first-step gates — they are cheap to check and each one has cost a full run before:
Check these at the first step (~20–40 minutes in):
• R3 on → pearson ≈ 0.999 in the log. Clearly below that means the routing replay is
misaligned and the gradients are already garbage — stop and investigate.
• lr is not 0 (fully-async + cosine collapses to 0; this runner forces constant, but
it is worth confirming).
• First-step reward / num_turns are non-zero — all zeros usually means the environment
images cannot be pulled, not a model problem.kill a live run to make room — ask the user to stop it themselveslib/site.env, or src/ to make a check passexp_name as the run's real name (Step 3)logs/<exp_name>.log as the training log without checking — that is
the old scripts' path; the config decides, and it is usually under
harbor_trials/. Read the runner's own train log: line instead of composing oneharbor_trials/, or checkpoints; a dangling
symlink under a run dir is enough to break the webui run list© LegoX, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .claude/plugins/rl-plugin/skills/run of LegoX/Lego-RL.
Open the folder on GitHubat commit 7c30234
Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Run this skillLegoX/Lego-RL | 108 | — | ~3k | Automated safety check: Pass | Apache-2.0 | |
| Shipping and Launch Checklistaddyosmani/agent-skills | 103k | 1 repos | ~2.8k | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 275k | — | ~2.2k | Automated safety check: Pass | MIT | |
| Evalalirezarezvani/claude-skills | 28k | 1 repos | ~618 | Automated safety check: Pass | MIT | |
| Eval Harnessaffaan-m/ECC | 275k | 1 repos | ~1.7k | Automated safety check: Pass | MIT | |
| Eval-Driven Development Harnessaffaan-m/ECC | 275k | — | ~1.5k | Automated safety check: Pass | MIT |
addyosmani/agent-skills
Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.
affaan-m/ECC
Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…
alirezarezvani/claude-skills
Evaluate and rank agent results by metric or LLM judge for an AgentHub session.
affaan-m/ECC
Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi
affaan-m/ECC
Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.
Orchestra-Research/AI-Research-SKILLs
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
LegoX/Lego-RL
Compose, edit, refactor, and validate Lego-RL train/eval/infer .env configs and reusable scripts/templates modules.
LegoX/Lego-RL
Preflight a Lego-RL config: answer "is it safe to launch this run right now?".
LegoX/Lego-RL
Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths.
LegoX/Lego-RL
Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy.
LegoX/Lego-RL
Guided install / scale-out of a sandbox Kubernetes cluster for the Lego-RL k8s backend (kubeadm 1.32 + containerd + flannel + ImageVolume, optionally nydus / a shared registry / an isolated dockerd).
LegoX/Lego-RL
One-to-one Codex counterpart for Claude /rl:check. An agent skill from LegoX/Lego-RL.
Preflight and launch a Lego-RL run (train, eval or infer). An agent skill from LegoX/Lego-RL. Run is an agent skill from LegoX/Lego-RL. Preflight and launch a Lego-RL run (train, eval or infer).
Run fits situations like: launch this config; run harbor training; kick off training; get this config running.
Run `npx skills add LegoX/Lego-RL --skill run -a claude-code`. Or copy the skill folder (.claude/plugins/rl-plugin/skills/run in LegoX/Lego-RL) into .claude/skills/run in your project. Claude Code loads it when a task matches its description.
Run `npx skills add LegoX/Lego-RL --skill run -a codex`. Or copy the skill folder (.claude/plugins/rl-plugin/skills/run in LegoX/Lego-RL) into .agents/skills/run in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LegoX/Lego-RL --skill run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run, .gemini/skills/run, .github/skills/run and .opencode/skills/run in your project.
Going by SKILL.md and its folder, Run needs the command-line tools its instructions call (bash, curl and kind).
SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Run is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Run: Shipping and Launch Checklist (addyosmani/agent-skills, 103k stars), Eval Harness (affaan-m/ECC, 275k stars), Eval (alirezarezvani/claude-skills, 28k stars) and Eval Harness (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
LegoX (a GitHub organization) maintains it in LegoX/Lego-RL, which has 108 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 8, 2026.
Source: LegoX/Lego-RL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.