Agent skill

Run

by LegoX in LegoX/Lego-RL

Preflight and launch a Lego-RL run (train, eval or infer). An agent skill from LegoX/Lego-RL.

Apache-2.0Auto-check passed

Install Run

skills CLI
$ npx skills add LegoX/Lego-RL --skill run -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LegoX/Lego-RL run --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/run .claude/skills/run && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
run
GitHub stars
108
Token cost
~3k tokens
SKILL.md length
1,166 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

Preflight and launch a Lego-RL run (train, eval or infer). An agent skill from LegoX/Lego-RL.

  • Works in 8 steps: Orient → Refuse if a run is already in flight → Preflight via /rl:check → …
  • Launch this config
  • SKILL.md covers Step 0 — Orient, Step 1 — Refuse if a run is…, Step 2 — Preflight via /rl:check and Step 3 — Show the parameters…, plus 5 more sections
  • Calls bash, curl and kind

What it does

Run is an agent skill from LegoX/Lego-RL. Preflight and launch a Lego-RL run (train, eval or infer). Runs /rl:check internally and refuses on any blocking failure, shows the fully resolved run parameters and waits for explicit confirmation, then launches the runner in the background by default — these runs take hours, so a foreground default would hold the session hostage. For multi-node train/infer it prints the exact per-node command instead of SSH-ing anywhere. Before launching it states where the run's log will actually land and whether the running…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Lego-RL: Harness-Native Reinforcement Learning for Coding Agents. The licence is Apache-2.0.

When your agent uses it

  • Launch this config
  • Run harbor training
  • Kick off training
  • Get this config running

Example prompts

  • “launch this config”
  • “run harbor training”
  • “start the eval”
  • “/run”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Orient
  2. Refuse if a run is already in flight
  3. Preflight via /rl:check
  4. Show the parameters and confirm (mandatory, never skip)
  5. Decide launch mode
  6. Multi-node: print the per-node command, do not SSH
  7. Launch
  8. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 7c30234. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash
    • curl
    • kind

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use curl, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Run loads about 3k tokens when it runs. Until then it costs about 202 tokens; SKILL.md has 1,166 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~202
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LegoX/Lego-RL at commit 7c30234, republished under its Apache-2.0 licence (© LegoX). 1,166 words, ~3,013 tokens.

Download SKILL.mdSave it as .claude/skills/run/SKILL.md (or your agent's skills folder).
name
run
description
Preflight and launch a Lego-RL run (train, eval or infer). Runs /rl:check internally and refuses on any blocking failure, shows the fully resolved run parameters and waits for explicit confirmation, then launches the runner in the background by default — these runs take hours, so a foreground default would hold the session hostage. For multi-node train/infer it prints the exact per-node command instead of SSH-ing anywhere. Before launching it states where the run's log will actually land and whether the running dashboard can see it, asking when it cannot. After launch it surfaces the real exp_name, the PIDs, the resolved log paths and the first-step gates worth watching. Triggers on "launch this config", "run harbor training", "start the eval", "kick off training", "get this config running".

/rl:run — launch one Lego-RL run

Preflight, review the parameters, confirm, then launch in the background. Refuses if any check blocks, or if a run is already live on this box.

Step 0 — Orient

Same as /rl:check Step 0: find the repo root, resolve the config (path / bare name / ask when absent or ambiguous), infer the kind (train | eval | infer). Never guess which config the user meant — a wrong guess here burns a cluster for hours.

Step 1 — Refuse if a run is already in flight

bash
bash scripts/lib/live_probe.sh <kind> 2>&1 | grep -E '^WARN +job:'

If any job:trainer or job:runner line is alive, abort:

A run is already in flight; refusing to start another:
  pids:  <pid + etime, from ps -p <pid> -o pid=,comm=,etime=>
  log:   <most recent *.log under logs/>
Let it finish, or stop it yourself (kill -INT <pid>), then re-run /rl:run.

A bare job:vllm with no runner above it is a foreign serving job, not ours — that is Step 2's business (it blocks on GPU ownership), not an abort message about "our" run.

Step 2 — Preflight via /rl:check

Run the check skill's logic on this config (invoke that skill, or inline its Steps 1–3). If the verdict is SAFE TO RUN: ❌ NO, abort with the same consolidated report, prefixed:

Preflight failed — fix the items below before /rl:run.

Do not invent values, skip a check, or pass a force flag; there is no force flag. If the user says "just launch it anyway", explain exactly which check blocks and what it costs to ignore it (each ✗ FATAL maps to a real past incident, written up in the Troubleshooting section of the docs site), then ask them to fix the config. Configs are user-owned; this skill is a launcher, not an editor.

Step 3 — Show the parameters and confirm (mandatory, never skip)

PREFLIGHT_ONLY=1 already printed the run configuration (<kind>) block in Step 2. Show that block — verbatim — and then draw attention to the handful of fields that decide whether the run is worth its hours. Keep this short; the block above already has the detail.

Worth double-checking before this run:
  • data     <train/val index, or dataset, or shard>   ← is this the one you meant?
  • scale    <N nodes × 8 gpu>, <trials/step>, epochs=<N>  ← roughly <estimate> hours/step
  • resume   <RESUME_FROM, or "auto: picks up the newest ckpt in the exp dir">
  • naming   project=<...>  exp=<...>
  • logs     <the real TRAIN_LOG path>
             dashboard <visible / not visible: it is serving <log-dir>>
  • first step  <R3 on → pearson ≈ 0.999; otherwise routing misalignment silently
                destroys the gradients>

Where the log lands, and whether the board will see it — this line is not cosmetic. scripts/templates/verl/common.env derives TRAIN_LOG=${HARBOR_LOG_DIR}/${TRAINER_EXPERIMENT_NAME}.log, and a real config overrides HARBOR_LOG_DIR to a per-experiment directory under the shared trials root. Only the template default is <repo>/logs. The dashboard globs its --log-dir one level, no recursion (webui/server.py) — so a config that overrides HARBOR_LOG_DIR produces a run that trains fine and is invisible on the board for its whole life unless something links it in. Resolve both sides before asking for the yes:

bash
bash scripts/train/train.sh --dry-run <config> 2>&1 | grep -E 'train log|vLLM log'
pgrep -af 'server\.py.*--log-dir'      # which dirs a board actually serves

If dirname(TRAIN_LOG) is not one of the served dirs — and "no board running at all" counts as not served — say so and ask. Do not launch silently, and do not pick for them:

This config writes its training log to
  <TRAIN_LOG>
but the dashboard (pid=<P>, port=<P>) is serving
  <served log-dir>
One-level glob, no match → this run will never appear on the board
(training itself is unaffected).

  [1] symlink it into <served-dir>/<name>.log right after launch
      (no dashboard restart — recommended)
  [2] leave it, watch wandb only
  [3] stop and change HARBOR_LOG_DIR in the config first
Which one?

Option 1 is one ln -s in Step 6, after the real exp_name is known. It is the only choice needing neither a restart nor a config edit, and the symlink name becomes the run's id in the UI — the analysis panels get their exp-dir mapping from the log contents, so any readable name works. Never move the log (the runner holds it open through tee), and never leave a dangling symlink under a log dir — that breaks /api/runs for every run.

Then ask, in one line:

Launch with this configuration? [Y]/[N]  (N = edit the config first)

Never launch without an explicit yes. If the user says no, stop cleanly and tell them to edit the config, then re-run /rl:run.

exp_name gotcha — unless the config pins EXP_NAME, train.sh derives it as harbor-<scaffold>-<EXP_TAG>-<timestamp> at launch time, so the name in the preflight output is not the name the run will have. Never report the preflight-time name as final; read the real one back in Step 6. If the user needs a stable name (resuming, comparing runs), suggest pinning EXP_NAME= in the config before launching.

Step 4 — Decide launch mode

Default = background. These runs last hours; foreground holds the session.

Launch mode? [B]ackground (default: nohup setsid, logs to logs/launch_<ts>.log)
             [F]oreground (holds this session until the run ends)

Accept B / F / <enter> (= background).

Step 5 — Multi-node: print the per-node command, do not SSH

Read topology (train) or vllm topo (infer) out of the parameter block.

  • train, NNODES > 1 — ray_bringup.sh branches on local IP: the node whose IP equals MASTER_ADDR (default: the launching node's first IP) becomes the ray head and drives training; every other node joins and exits. So the same command must be run on each of the NNODES nodes.
  • infer, VLLM_NNODES > 1 — same shape, keyed on VLLM_HEAD_HOST; rank-0 serves the API, the others run --headless DP shards.
  • eval — always single-node.

Print the command block and tell the user to run it on the other nodes themselves. Never SSH to another node, and never assume the workers are up just because the head started.

Multi-node: run the command below once on every node (this machine is already covered
by the launch just performed)
  nodes <ip2>, <ip3>, ...:
    cd <repo_root> && MASTER_ADDR=<head_ip> nohup setsid bash scripts/train/train.sh <config> > logs/launch_<ts>_$(hostname -s).log 2>&1 &
The head node waits for all <NNODES> nodes to join before training starts; a worker that
never came up shows as the head spinning in its ray status poll.
Show full SKILL.md (458 more words)Show less

Step 6 — Launch

Timestamp: TS=$(date -u +%Y%m%d-%H%M%S). Log: logs/launch_${TS}.log (mkdir -p logs first).

Background (default)
bash
cd <repo_root> && nohup setsid bash scripts/<kind>/<kind>.sh <config> > "logs/launch_${TS}.log" 2>&1 < /dev/null &
PARENT_PID=$!
disown

setsid matters: it detaches the run from this session so a disconnect does not take the training down with it.

Then, without polling in a tight loop:

  1. Wait up to 60s (say, 5s sleeps) for the launch log to grow past 0 bytes.

  2. Read its first ~120 lines and pull out:

    • the real exp= value from the run configuration block,
    • the preflight verdict line (it runs again inside the real launch),
    • any early [FATAL].
  3. Once they appear, capture the PIDs: pgrep -f 'scripts/<kind>/<kind>.sh' and, for train, pgrep -f 'fully_async_main|main_ppo' (the trainer takes a few minutes to show up — report it as pending rather than waiting for it).

  4. If Step 3 chose [1], link the log into the served dir now that exp_name is real. Verify with the API rather than assuming — the run only shows up once it has training steps, so an empty steps right after launch is expected:

    bash
    ln -sfn "<TRAIN_LOG>" "<served-dir>/<readable-name>.log"
    curl -s -m 300 "http://127.0.0.1:<P>/api/runs" | grep -c '<readable-name>'

    For a multi-node run, link only the ray-head node's log — that is the one carrying the trainer output; the worker exp dirs hold just *_train_gpu_wandb.log, and the analysis panels recover the sibling exp dirs on their own.

If the log stays empty for 60s, or the runner exits non-zero within 60s, print the tail of the launch log and stop. Do not retry automatically and do not "fix" the config to get past it.

Foreground (only if chosen)
bash
cd <repo_root> && bash scripts/<kind>/<kind>.sh <config> 2>&1 | tee "logs/launch_${TS}.log"

Stream it; on exit report the exit code and the log tail.

Step 7 — Report

Tight summary, ≤ 15 lines:

Launched (background)
  kind / config:  <kind> · <config path relative to the repo>
  exp_name:       <the real name read back from the launch log; if not there yet,
                   "starting — check the log in 30s">
  launch log:     logs/launch_<TS>.log
  train log:      <absolute TRAIN_LOG path, read from the launch log's "train log:" line>
  vllm log:       <absolute VLLM_LOG path, same source>
  trials:         <HARBOR_TRIALS_DIR>
  dashboard:      <URL, run name = <readable-name> / not visible (user chose [2])>
  pids:           runner=<pid>  trainer=<pid or "starting">
  stage:          vLLM loading weights + cuda-graph capture; ~10–20 min before step 1

Watch:
  tail -F logs/launch_<TS>.log
  nvidia-smi
  /rl:status            # once it is running, check that the first step is healthy
Stop:
  kill -INT <runner_pid>    # let it wind down; do not kill -9

For train, close with the first-step gates — they are cheap to check and each one has cost a full run before:

Check these at the first step (~20–40 minutes in):
  • R3 on → pearson ≈ 0.999 in the log. Clearly below that means the routing replay is
    misaligned and the gradients are already garbage — stop and investigate.
  • lr is not 0 (fully-async + cosine collapses to 0; this runner forces constant, but
    it is worth confirming).
  • First-step reward / num_turns are non-zero — all zeros usually means the environment
    images cannot be pulled, not a model problem.

What this skill must NOT do

  • never force past a blocking preflight, and never add a flag to bypass one
  • never kill a live run to make room — ask the user to stop it themselves
  • never edit the config, lib/site.env, or src/ to make a check pass
  • never SSH to another node, or launch anything on a node other than this one
  • never run in the foreground without asking — it locks the session for hours
  • never report the preflight-time exp_name as the run's real name (Step 3)
  • never print logs/<exp_name>.log as the training log without checking — that is the old scripts' path; the config decides, and it is usually under harbor_trials/. Read the runner's own train log: line instead of composing one
  • never launch without telling the user whether the dashboard will see this run
  • never delete or move logs, harbor_trials/, or checkpoints; a dangling symlink under a run dir is enough to break the webui run list

© LegoX, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/plugins/rl-plugin/skills/run of LegoX/Lego-RL.

Open the folder on GitHubat commit 7c30234

Compare with similar skills

Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Run compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Run this skillLegoX/Lego-RL108—~3kAutomated safety check: PassApache-2.0
Shipping and Launch Checklistaddyosmani/agent-skills103k1 repos~2.8kAutomated safety check: PassMIT
Eval Harnessaffaan-m/ECC275k—~2.2kAutomated safety check: PassMIT
Evalalirezarezvani/claude-skills28k1 repos~618Automated safety check: PassMIT
Eval Harnessaffaan-m/ECC275k1 repos~1.7kAutomated safety check: PassMIT
Eval-Driven Development Harnessaffaan-m/ECC275k—~1.5kAutomated safety check: PassMIT

Similar skills

  • Shipping and Launch Checklist

    addyosmani/agent-skills

    Prepares a production launch with a pre-launch checklist, monitoring, a staged rollout and a rollback plan so every release is reversible and observable.

    103k GitHub starsUsed in 1 repo~2.8k tokens
    DevOps & CloudAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) framework for AI coding sessions — define capability and regression evals before coding, grade with code-based, model-based, rule, or human graders, and track pass@k…

    275k GitHub stars~2.2k tokensUpdated 3 days ago
    AI & LLM EngineeringAuto-check passed
  • Eval

    alirezarezvani/claude-skills

    Evaluate and rank agent results by metric or LLM judge for an AgentHub session.

    28k GitHub starsUsed in 1 repo~618 tokens
    AI & LLM EngineeringAuto-check passed
  • Eval Harness

    affaan-m/ECC

    Eval-driven development (EDD) ilkelerini uygulayan Claude Code oturumları için formal değerlendirme çerçevesi

    275k GitHub starsUsed in 1 repo~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Sets up eval-driven development for Claude Code workflows: capability and regression evals, three grader types and pass@k reliability metrics.

    275k GitHub stars~1.5k tokensUpdated 3 days ago
    Agent WorkflowsAuto-check passed
  • Ray Train Distributed Training

    Orchestra-Research/AI-Research-SKILLs

    Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

    13k GitHub starsUsed in 3 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed

More from LegoX/Lego-RL

All 11 skills in this repo
  • Lego Rl Config

    LegoX/Lego-RL

    Compose, edit, refactor, and validate Lego-RL train/eval/infer .env configs and reusable scripts/templates modules.

    108 GitHub stars~2.1k tokensUpdated today
    Auto-check: notes
  • Check

    LegoX/Lego-RL

    Preflight a Lego-RL config: answer "is it safe to launch this run right now?".

    108 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Dashboard

    LegoX/Lego-RL

    Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths.

    108 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Status

    LegoX/Lego-RL

    Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy.

    108 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • K8s Sandbox Install

    LegoX/Lego-RL

    Guided install / scale-out of a sandbox Kubernetes cluster for the Lego-RL k8s backend (kubeadm 1.32 + containerd + flannel + ImageVolume, optionally nydus / a shared registry / an isolated dockerd).

    108 GitHub stars~2.9k tokensUpdated today
    Auto-check: warnings
  • Rl Check

    LegoX/Lego-RL

    One-to-one Codex counterpart for Claude /rl:check. An agent skill from LegoX/Lego-RL.

    108 GitHub stars~338 tokensUpdated today
    Auto-check passed

Questions about Run

What does Run do?

Preflight and launch a Lego-RL run (train, eval or infer). An agent skill from LegoX/Lego-RL. Run is an agent skill from LegoX/Lego-RL. Preflight and launch a Lego-RL run (train, eval or infer).

When should I use Run?

Run fits situations like: launch this config; run harbor training; kick off training; get this config running.

How do I install Run in Claude Code?

Run `npx skills add LegoX/Lego-RL --skill run -a claude-code`. Or copy the skill folder (.claude/plugins/rl-plugin/skills/run in LegoX/Lego-RL) into .claude/skills/run in your project. Claude Code loads it when a task matches its description.

How do I install Run in Codex?

Run `npx skills add LegoX/Lego-RL --skill run -a codex`. Or copy the skill folder (.claude/plugins/rl-plugin/skills/run in LegoX/Lego-RL) into .agents/skills/run in your project. Codex loads it when a task matches its description.

Can I use Run in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LegoX/Lego-RL --skill run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/run, .gemini/skills/run, .github/skills/run and .opencode/skills/run in your project.

What does Run need to run?

Going by SKILL.md and its folder, Run needs the command-line tools its instructions call (bash, curl and kind).

Does Run access the network?

SKILL.md contains no URLs. Its commands use curl, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Run safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Run use?

Run is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Run use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Run?

Skills that share tags, products or a category with Run: Shipping and Launch Checklist (addyosmani/agent-skills, 103k stars), Eval Harness (affaan-m/ECC, 275k stars), Eval (alirezarezvani/claude-skills, 28k stars) and Eval Harness (affaan-m/ECC, 275k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Run?

LegoX (a GitHub organization) maintains it in LegoX/Lego-RL, which has 108 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 8, 2026.

Source: LegoX/Lego-RL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.