Agent skill

Status

by LegoX in LegoX/Lego-RL

Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy.

Apache-2.0Auto-check passed

Install Status

skills CLI
$ npx skills add LegoX/Lego-RL --skill status -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install LegoX/Lego-RL status --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/LegoX/Lego-RL.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/plugins/rl-plugin/skills/status .claude/skills/status && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
status
GitHub stars
108
Token cost
~3k tokens
SKILL.md length
1,272 words
Files
1
Skills in repo
11
Repo updated
First seen
Licence
Apache-2.0

At a glance

Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy.

  • Works in 5 steps: Which run? → How far has it got? → Are the numbers healthy? → …
  • Hows the run doing
  • SKILL.md covers Step 1 — Which run?, Step 2 — How far has it got?, Step 3 — Are the numbers… and Step 4 — Match against known…, plus 2 more sections
  • Calls bash

What it does

Status is an agent skill from LegoX/Lego-RL. Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy. Reads the process table, the run log's metric lines and the trials directory, then checks the metrics against this cluster's known failure signatures — R3 pearson collapse, lr=0, grad starvation, envsetupfailed avalanches, val fake-zeros, no-tool-call collapse, and for SAO/critic runs critic starvation and the all-negative-batch collapse — and says which one matches…

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Lego-RL: Harness-Native Reinforcement Learning for Coding Agents. The licence is Apache-2.0.

When your agent uses it

  • Hows the run doing
  • What step is it on
  • Is reward going up
  • Is this run broken

Example prompts

  • “s the run doing”
  • “what step is it on”
  • “is reward going up”
  • “/status”

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Which run?
  2. How far has it got?
  3. Are the numbers healthy?
  4. Match against known failure signatures
  5. Report

What it can do on your machine

Read from SKILL.md and the folder at commit 7c30234. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Status loads about 3k tokens when it runs. Until then it costs about 181 tokens; SKILL.md has 1,272 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~181
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from LegoX/Lego-RL at commit 7c30234, republished under its Apache-2.0 licence (© LegoX). 1,272 words, ~3,003 tokens.

Download SKILL.mdSave it as .claude/skills/status/SKILL.md (or your agent's skills folder).
name
status
description
Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy. Reads the process table, the run log's metric lines and the trials directory, then checks the metrics against this cluster's known failure signatures — R3 pearson collapse, lr=0, grad starvation, env_setup_failed avalanches, val fake-zeros, no-tool-call collapse, and for SAO/critic runs critic starvation and the all-negative-batch collapse — and says which one matches. Read-only, local host only: never kills, never restarts, never edits. Triggers on "how's the run doing", "what step is it on", "is reward going up", "is this run broken", "diagnose the training run".

/rl:status — is the live run healthy?

Read-only diagnosis of a run in flight. Answers three things in order: which run · how far · is it sick. Never kills, restarts, cleans or edits anything — a wrong intervention here costs more than a slow answer.

Step 1 — Which run?

bash
bash scripts/lib/live_probe.sh train 2>&1 | grep -E '^(WARN|OK) +(job|gpu):'
ls -t logs/*.log | head -5

Identify the live run from the job:trainer / job:runner lines, then map it to its log.

Do not assume the log is under logs/. scripts/templates/verl/common.env derives TRAIN_LOG=${HARBOR_LOG_DIR}/${TRAINER_EXPERIMENT_NAME}.log, and a real config overrides HARBOR_LOG_DIR to a per-experiment directory under the shared trials root; only the template default lands in <repo>/logs. Get the real path from the runner itself, in this order:

bash
# 1. the runner printed it at startup (works even for a run launched by hand)
grep -hoE 'train log: +\S+' logs/launch_*.log *.out 2>/dev/null | tail -3
# 2. or resolve it from the config without launching anything
bash scripts/train/train.sh --dry-run <config> 2>&1 | grep -E 'train log|vLLM log|trials'
# 3. or find what is actually being written right now
find "$(dirname "$HARBOR_TRIALS_DIR")" -mindepth 3 -maxdepth 3 -type d -name logs \
     -mmin -30 2>/dev/null | head   # or point it at your trials root

A run launched by hand as nohup bash scripts/train/train.sh <config> > foo.out leaves foo.out wherever the launcher's cwd was — usually the repo root, not logs/. It holds the launch summary plus the same teed stream, so it is a superset of TRAIN_LOG and equally good to read; the train log: line near its top is the fastest way to recover the canonical path. What it is not is a file the dashboard can see, since it is outside any served log dir.

Because the repo lives on shared storage, that .out is visible from every box while the process is not: on a multi-node run only the ray-head node has the train.sh / tee / trainer processes. Seeing a growing log with no matching pid here means you are on the wrong node — not that the run died. Check stat -c %Y on the log before concluding anything from an empty pgrep.

Multi-node: each node evaluates EXP_NAME=…$(date …) separately, so one launch produces NNODES exp dirs whose timestamps differ by seconds. Only the ray-head dir holds the trainer log; the others hold just *_train_gpu_wandb.log. Diagnose from the head's, but remember trial counts must be summed across all sibling dirs.

If nothing is alive, say so and offer the last finished run instead; make it explicit in the report which of the two you are describing. If several runs are alive, list them and ask which one — do not merge metrics from two runs.

Step 2 — How far has it got?

bash
LOG=<TRAIN_LOG resolved in Step 1>                # NOT assumed to be logs/<exp>.log
grep -oE 'step:[0-9]+ ' "$LOG" | tail -1          # latest step
grep -cE ' step:[0-9]+ - training/global_step' "$LOG"
ls -t harbor_trials/<project>/<exp_name> 2>/dev/null | head -3
tail -40 "$LOG"

Report: latest step, wall-clock since launch (ps -p <pid> -o etime=), average minutes/step, and whether the tail is still moving (compare stat -c %Y "$LOG" against now). A log that has not been written to in >30 min while the process is alive is itself the finding — that is the deadlock shape, not a slow step.

Step 3 — Are the numbers healthy?

Metrics live on the step lines as key:value pairs. Pull the latest step line and read the keys below (these names are exact — they come from the real logs):

bash
grep -E ' step:[0-9]+ - training/global_step' "$LOG" | tail -1 \
  | grep -oE '(training/rollout_actor_probs_pearson_corr|actor/(lr|grad_norm|kl_coef|pg_clipfrac)|actor/rollout_corr/(kl|rollout_is_eff_sample_size|rollout_is_ratio_fraction_low)|critic/(rewards/mean|advantages/mean|vf_explained_var|vf_loss|grad_norm|lr)|num_turns/mean|trajectory_filter/[a-z_/]+|response_length/(mean|clip_ratio)):[0-9.e+-]+'

Tell a SAO / critic run apart first: its config block says gae with a critic line, and the log carries critic/vf_explained_var. On such a run training/rollout_actor_probs_pearson_corr and actor/entropy are absent by design (bypass mode: old_log_probs == rollout_log_probs, so pearson would be 1.0 by construction) — their absence is not the R3 signature. Read the actor/rollout_corr/* keys instead.

Metric keyHealthyWhat a bad value means
training/rollout_actor_probs_pearson_corr≈ 0.999 (≥ 0.99)R3 routing replay is misaligned — training on corrupted logprobs. The single most important gate; a run below this is already wasted.
actor/lr= the configured lr0 → the fully-async + cosine + total_training_steps=-1 bug; the model is frozen. Runner forces constant, so a 0 here means something overrode it.
actor/grad_normsame order as prior runs (~0.2–0.5)~0.03 with very long responses = gradient starvation from token dilution, not a bug to fix mid-run.
critic/rewards/meannon-zero, trending upFlat 0 from step 1 = infrastructure, not the model — go to the filter reasons below before touching hyperparameters.
num_turns/meantens of turnsCollapsing toward ~1 with reward dropping = the model stopped emitting tool calls and just ends the episode; a real training pathology, not infra.
trajectory_filter/reason/env_setup_failed~0Non-trivial count = pods cannot start: image unpullable, registry down, or a node missing its insecure-registry trust.
trajectory_filter/reason/timeoutsmall fractionA large share means the agent budget is too tight for these tasks, or env exec is stalling.
trajectory_filter/invalid_ratio< ~0.1High = most of the batch is being dropped; the effective batch is far smaller than configured.
response_length/clip_ratiolowHigh = responses hitting the window; the tail is being truncated.
val-core/…, val-aux/num_turns/…non-zero at test_freq stepsAll-zero val while train reward is fine = the val split's images are unpullable, not a model regression.
critic/vf_explained_var (SAO)leaves <0 within ~20 steps, then 0.2–0.5Flat ≤ 0.3 for 50+ steps = the critic never converged; with critic/grad_norm far above CRITIC_GRAD_CLIP that is critic starvation (every update clipped down). Huge negatives on a step whose critic/returns/min ≈ max are a degenerate batch, ignore that step.
critic/grad_norm (SAO)median ~10 on 30B–35B, spikes to 30–70Alarm only on three consecutive steps > 30; a single spike (even 200+, e.g. an empty batch after sandboxes vanished) is not instability.
critic/advantages/mean (SAO)≈ 0 with whitening onDrifting negative for consecutive steps with whitening off = all-negative batches; the precursor of the think-spam collapse.
actor/rollout_corr/rollout_is_eff_sample_size (SAO)≥ 0.99Well below = DIS is zeroing many tokens (staleness or backend mismatch); with rollout_is_ratio_fraction_low at an exact multiple of 1/batch the DIS mirror is missing and the run trains sequence-TIS.
actor/pg_clipfrac (SAO)absentPresent on a bypass-mode run = the actor is not in bypass_mode; DIS never reached the loss.

Also worth a line each when present: fully_async/processing_time/tp99 (long tail), fully_async/count/dropped_stale_samples (staleness pressure), rollout_corr/kl.

Show full SKILL.md (390 more words)Show less

Step 4 — Match against known failure signatures

Only claim a signature when its specific evidence is present. Say "no known signature matched" rather than forcing a match — a wrong diagnosis here sends the user chasing the wrong layer for hours.

SignatureEvidence that must be present
R3 misalignmentpearson well below 0.99 on recent steps
frozen modelactor/lr:0
grad starvationactor/grad_norm an order below the run's own earlier steps, alongside very long response_length/mean
env avalanchetrajectory_filter/reason/env_setup_failed climbing across steps; reward down in step
val fake-zeroval metrics 0 while critic/rewards/mean is healthy
no-tool-call collapsenum_turns/mean falling toward 1 over consecutive steps + reward falling; filter reasons normal
critic starvation (SAO)critic/vf_explained_var flat ≤ 0.3 for 50+ steps while critic/grad_norm sits well above the configured clip; val flat. Fix on the next run: CRITIC_GRAD_CLIP=10, a warm CRITIC_MODEL_PATH, CRITIC_WARMUP=20
all-negative-batch collapse (SAO)critic/advantages/mean negative on 3+ consecutive steps (whitening off) followed by num_turns/mean rising while reward falls — a reward-neutral tool (e.g. think) is being relatively reinforced. GAE_WHITEN_ADVANTAGES=True on the next run; roll back to before the drift
DIS not reaching the actor (SAO)actor/pg_clipfrac present, or rollout_is_ratio_fraction_low landing on exact multiples of 1/batch — the policy-loss mirror in hydra_args.sh was removed; the run is not SAO
deadlock / stallprocess alive, log mtime old, no new step line; check whether the tail sits in val or in a rollout wait
step slowdownminutes/step up sharply — compare the node/replica counts in the run's own config block before blaming the tasks

For anything that points off-box (registry, kyverno, node disk, image pulls), report the symptom and stop. This skill does not SSH, does not touch the cluster, and must not assert a cluster-side cause it cannot see from here — phrase it as "the symptom points at X; confirm on <node>", and let the user decide.

Step 5 — Report

Keep metric keys verbatim so they can be grepped.

## harbor status — <exp_name>

**<🟢 healthy | 🟡 at risk | 🔴 recommend stopping>** — <one-line conclusion>

  stage    step <N> (<epoch>) · running <etime> · ~<M> min/step · log last written <X> min ago
  procs    runner=<pid>  trainer=<pid>  GPUs in use: <n>
  reward   critic/rewards/mean=<v> (last <k> steps: <trend>)
  grads    actor/grad_norm=<v>   actor/lr=<v>   pearson=<v | n/a (bypass mode)>
  critic   vf_explained_var=<v>  grad_norm=<v>  ESS=<v>        (SAO runs only)
  traj     num_turns/mean=<v>  invalid_ratio=<v>  env_setup_failed=<v> timeout=<v>
  val      <value from the most recent val, or "test_freq not reached yet">

**Diagnosis**
<the matched signature + its supporting evidence; otherwise "no known signature matched">

**Recommendations**
1. <at most 3, cheapest first; irreversible actions such as stopping a run are always
   phrased as recommendations for the user to carry out>

Never end with an action you already took — this skill takes none.

Guardrails

  • read-only: no kill, no ray stop, no restart, no config edit, no log deletion or rotation (a dangling symlink under a run dir breaks the webui)
  • local host only: no SSH, no kubectl mutation
  • never merge two runs' metrics into one report
  • never state a cause you did not read out of the log or the process table
  • when the evidence is thin, say the evidence is thin

© LegoX, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/plugins/rl-plugin/skills/status of LegoX/Lego-RL.

Open the folder on GitHubat commit 7c30234

Compare with similar skills

Status next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Status compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Status this skillLegoX/Lego-RL108—~3kAutomated safety check: PassApache-2.0
Flightsasgeirtj/system_prompts_leaks69k—~2.1kAutomated safety check: PassCC0-1.0
Flightscodewhale-hq/Codewhale41k—~258Automated safety check: PassMIT
Diagnose Gatewayopenclaw/openclaw392k—~670Automated safety check: PassMIT
Diagnosegithub/awesome-copilot40k1 repos~1kAutomated safety check: PassMIT
Diagnose Why Work Stoppedpaperclipai/paperclip99k—~2.8kAutomated safety check: PassMIT

Similar skills

  • Flights

    asgeirtj/system_prompts_leaks

    Research, book, or manage flights, including check-in. An agent skill from asgeirtj/system_prompts_leaks.

    69k GitHub stars~2.1k tokensUpdated today
    Auto-check passed
  • Flights

    codewhale-hq/Codewhale

    Track flights and look up status and schedules. An agent skill from codewhale-hq/Codewhale.

    41k GitHub stars~258 tokensUpdated today
    Auto-check passed
  • Diagnose Gateway

    openclaw/openclaw

    Diagnose Gateway, config, secrets, channels, and port failures with read-only one-liners.

    392k GitHub stars~670 tokensUpdated today
    Auto-check passed
  • Diagnose

    github/awesome-copilot

    Official

    Perform a systematic diagnostic scan of an AI workflow across 5 quality dimensions — prompt quality, context efficiency, tool health, architecture fitness, and safety — producing a scored report…

    40k GitHub starsUsed in 1 repo~1k tokens
    Agent WorkflowsAuto-check passed
  • Diagnose Why Work Stopped

    paperclipai/paperclip

    Diagnose stalled, looping, or over-recovered Paperclip issue trees and propose a no-code product-rule plan.

    99k GitHub stars~2.8k tokensUpdated today
    Auto-check passed
  • Investigates a session where Superpowers went wrong, reads the transcripts on disk and produces an evidence-cited report, optionally prepared as a bug report for the maintainers.

    296k GitHub starsUsed in 3 repos~1.7k tokens
    Agent WorkflowsAuto-check passed

More from LegoX/Lego-RL

All 11 skills in this repo
  • Lego Rl Config

    LegoX/Lego-RL

    Compose, edit, refactor, and validate Lego-RL train/eval/infer .env configs and reusable scripts/templates modules.

    108 GitHub stars~2.1k tokensUpdated today
    Auto-check: notes
  • Check

    LegoX/Lego-RL

    Preflight a Lego-RL config: answer "is it safe to launch this run right now?".

    108 GitHub stars~2.9k tokensUpdated today
    Auto-check passed
  • Dashboard

    LegoX/Lego-RL

    Bring up the Lego-RL training dashboard (webui/) on whatever machine you are on, adapting to that box's layout instead of assuming this repo's paths.

    108 GitHub stars~4.2k tokensUpdated today
    Auto-check passed
  • Run

    LegoX/Lego-RL

    Preflight and launch a Lego-RL run (train, eval or infer). An agent skill from LegoX/Lego-RL.

    108 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • K8s Sandbox Install

    LegoX/Lego-RL

    Guided install / scale-out of a sandbox Kubernetes cluster for the Lego-RL k8s backend (kubeadm 1.32 + containerd + flannel + ImageVolume, optionally nydus / a shared registry / an isolated dockerd).

    108 GitHub stars~2.9k tokensUpdated today
    Auto-check: warnings
  • Rl Check

    LegoX/Lego-RL

    One-to-one Codex counterpart for Claude /rl:check. An agent skill from LegoX/Lego-RL.

    108 GitHub stars~338 tokensUpdated today
    Auto-check passed

Questions about Status

What does Status do?

Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy. Status is an agent skill from LegoX/Lego-RL. Diagnose a Lego-RL run that is already in flight (or just finished): which run is alive, how far it has got, and whether its numbers are healthy.

When should I use Status?

Status fits situations like: hows the run doing; what step is it on; is reward going up; is this run broken.

How do I install Status in Claude Code?

Run `npx skills add LegoX/Lego-RL --skill status -a claude-code`. Or copy the skill folder (.claude/plugins/rl-plugin/skills/status in LegoX/Lego-RL) into .claude/skills/status in your project. Claude Code loads it when a task matches its description.

How do I install Status in Codex?

Run `npx skills add LegoX/Lego-RL --skill status -a codex`. Or copy the skill folder (.claude/plugins/rl-plugin/skills/status in LegoX/Lego-RL) into .agents/skills/status in your project. Codex loads it when a task matches its description.

Can I use Status in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add LegoX/Lego-RL --skill status -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/status, .gemini/skills/status, .github/skills/status and .opencode/skills/status in your project.

What does Status need to run?

Going by SKILL.md and its folder, Status needs the command-line tools its instructions call (bash).

Does Status access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Status safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Status use?

Status is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Status use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Status?

Skills that share tags, products or a category with Status: Flights (asgeirtj/system_prompts_leaks, 69k stars), Flights (codewhale-hq/Codewhale, 41k stars), Diagnose Gateway (openclaw/openclaw, 392k stars) and Diagnose (github/awesome-copilot, 40k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Status?

LegoX (a GitHub organization) maintains it in LegoX/Lego-RL, which has 108 GitHub stars. The repository holds 11 skills in this directory. The repository was last updated on October 8, 2026.

Source: LegoX/Lego-RL on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.