Agent skill

Test Install

by RLinf in RLinf/RLinf

Test that requirements/install.sh works for an embodied model/env by building its venv and running the matching CI e2e test.

Apache-2.0Auto-check passedTesting & QA

Install Test Install

skills CLI
$ npx skills add RLinf/RLinf --skill test-install -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install RLinf/RLinf test-install --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/RLinf/RLinf.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/test-install .claude/skills/test-install && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
test-install
GitHub stars
5.4k
Token cost
~2.3k tokens
SKILL.md length
965 words
Files
2
Skills in repo
9
Repo updated
First seen
Licence
Apache-2.0

At a glance

Test that requirements/install.sh works for an embodied model/env by building its venv and running the matching CI e2e test.

  • Verify an install
  • SKILL.md covers Prerequisites, Run (agent path), Verifying a brand-new… and Gotchas, plus 1 more section
  • Runs Python scripts from its folder; calls python3, uv and git; reaches gh-proxy.com
  • Check a new model/env installs cleanly

What it does

Test Install is an agent skill from RLinf/RLinf. Test that requirements/install.sh works for an embodied model/env by building its venv and running the matching CI e2e test. Use when asked to test or verify an install, check a new model/env installs cleanly, run the e2e test for a model/env, confirm a venv works, or check that the model/checkpoint paths referenced by an e2e config actually exist on disk.

Its SKILL.md is about 2.3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 1 other file (for example `driver.py`).

It sits in Testing & QA, covering End-to-end testing. The repository describes itself as: RLinf: Reinforcement Learning Infrastructure for Embodied and Agentic AI. The licence is Apache-2.0.

When your agent uses it

  • Verify an install
  • Check a new model/env installs cleanly
  • Run the e2e test for a model/env
  • Confirm a venv works

Example prompts

  • “/test-install”

Requirements

  • Python 3
  • Docker

What it can do on your machine

Read from SKILL.md and the folder at commit c70606f. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python), which the agent can run.

    Shell commands in SKILL.md call:

    • python3
    • uv
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • gh-proxy.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Test Install loads about 2.3k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 965 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~2.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from RLinf/RLinf at commit c70606f, republished under its Apache-2.0 licence (© RLinf). 965 words, ~2,322 tokens.

Download SKILL.mdSave it as .claude/skills/test-install/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
test-install
description
Test that requirements/install.sh works for an embodied model/env by building its venv and running the matching CI e2e test. Use when asked to test or verify an install, check a new model/env installs cleanly, run the e2e test for a model/env, confirm a venv works, or check that the model/checkpoint paths referenced by an e2e config actually exist on disk.

Verify that requirements/install.sh actually works for an embodied model/env: build its venv, confirm the e2e config's model paths exist, then run the matching CI e2e test against that venv. The harness is .agents/skills/test-install/driver.py — it reads the install command, env vars, and test config straight out of .github/workflows/embodied-e2e-tests.yml, so it never drifts from CI. Drive everything through that script.

All paths below are relative to the repo root (the dir with requirements/install.sh).

Prerequisites

This runs on the embodied CI runner (or an equivalent box): NVIDIA GPUs, the shared /workspace/dataset/ tree (models, LIBERO, etc.), uv, and a python3 that has PyYAML. No apt-get needed — the driver only orchestrates. Quick check:

bash
nvidia-smi -L | head -1
ls -d /workspace/dataset >/dev/null && python3 -c "import yaml" && echo "env OK"

If python3 lacks PyYAML, the driver auto-reexecs under uv run --with pyyaml, so it works regardless.

Run (agent path)

The driver has seven subcommands. Start with the read-only ones (list, resolve, check-paths) — they're instant and tell you exactly what CI does before you spend an hour on an install.

bash
# What model/env combos does CI cover, and which configs do they run?
python3 .agents/skills/test-install/driver.py list

# Show the exact install command + env vars + test configs for one combo:
python3 .agents/skills/test-install/driver.py resolve gr00t_n1d6 maniskill_libero

# Do the model/checkpoint paths an e2e config needs actually exist on disk?
python3 .agents/skills/test-install/driver.py check-paths libero_spatial_ppo_gr00t_n1d6

# Same check across every e2e config at once (great pre-flight / PR check):
python3 .agents/skills/test-install/driver.py check-all

check-paths classifies every absolute path in the config: [model/input] (must exist — a missing one fails the run before training and returns exit 1), [output dir] (created by the run, may be missing), [path] (informational).

The full pipeline

run does install → check-paths (per config) → test, stopping a test whose required model paths are missing. Always preview with --dry-run first — it prints the exact shell (install line + CI test step) without executing:

bash
python3 .agents/skills/test-install/driver.py run gr00t_n1d6 maniskill_libero \
    --venv /workspace/test-venvs/gr00t_n1d6 --dry-run

Drop --dry-run to actually build and test. Or drive the two halves separately (faster iteration — install once, test many):

bash
# Preview the install (env vars + install.sh line, with --venv and --use-mirror
# injected). Drop --dry-run to build the venv; add --no-mirror to skip mirrors:
python3 .agents/skills/test-install/driver.py install gr00t_n1d6 maniskill_libero \
    --venv /workspace/test-venvs/gr00t_n1d6 --dry-run

# Run the matching e2e test against any venv (verbatim CI test step, venv path
# swapped in). This one is cheap — dummy SAC, 2 epochs — so run it for real:
python3 .agents/skills/test-install/driver.py test realworld_dummy_sac_cnn \
    --venv /opt/venv/openvla

A full model install (e.g. gr00t_n1d6) is heavy — it clones repos and builds flash-attn — so preview with --dry-run, then run it where you can afford the time. The dummy-SAC test above completes in a few minutes against the prebuilt /opt/venv/openvla.

A real test run spins up Ray, the env/rollout/policy workers, and trains for the config's (deliberately tiny) epoch count. The dummy-SAC smoke above finishes in a few minutes and writes TensorBoard output under the config's log_path (/workspace/results/<config>/tensorboard/) — that directory appearing with fresh files is your "it worked" signal.

Clean up the venv when you're done

A test venv is heavy (e.g. gr00t_n1d6 is ~600M) and is throwaway — it exists only to prove the install + e2e work. After the test finishes, rm -rf the --venv path you built unless the user asked to keep it for reuse:

bash
rm -rf /workspace/test-venvs/<model>

Don't touch the shared caches (/workspace/dataset/.uv, .uv_cache) or the prebuilt /opt/venv/* — only the per-test venv you created. Cleaning up keeps /workspace/test-venvs/ from accumulating stale multi-hundred-MB trees across runs. (Cleanup is for the venv only; the /workspace/results/<config>/ outputs are your proof the test ran — leave them or mention them.)

When there's no matching test

If a model/env has an install but no e2e job in the workflow, run installs and then tells you there's nothing to run — ask the user what to run rather than guessing. If you have a config name that isn't wired into CI, test <config> --runner run|run_async|run_offline runs it directly (pick the runner; run is the default).

Verifying a brand-new model/env (the common case)

When someone adds install_<model>_model() + an e2e job (see the add-install-docker-ci-e2e and install-check skills), confirm it end to end:

bash
# 1. CI parsed it correctly and the install command looks right:
python3 .agents/skills/test-install/driver.py resolve <model> <env>
# 2. The SFT checkpoint the e2e config points at is actually on this box:
python3 .agents/skills/test-install/driver.py check-paths <its_config>
# 3. Full build + test (preview with --dry-run, then drop it to run for real):
python3 .agents/skills/test-install/driver.py run <model> <env> \
    --venv /workspace/test-venvs/<model> --dry-run

If step 2 reports MISS [model/input], the install can be perfect and the e2e will still fail — the dataset/checkpoint just isn't staged on this runner. Surface that to the user; it's not an install bug.

Show full SKILL.md (413 more words)Show less

Gotchas

  • install/run need --venv <path> — it's injected as install.sh --venv. Use an absolute path (e.g. /workspace/test-venvs/<model>); a bare name lands relative to the repo root. The test step then sources <path>/bin/activate.
  • --venv reuses an existing venv if one is already at that path (install.sh validates the Python version and reuses it). For a truly clean install, point at a fresh path or rm -rf it first.
  • The driver runs CI's shell verbatim, including its export UV_PATH=/workspace/dataset/.uv etc. That's intentional — it reproduces CI exactly. It also means the install writes into the shared uv cache, same as CI.
  • --use-mirror is added to every install by default (faster downloads), even for CI jobs that don't list it. Pass --no-mirror to install/run to turn it off. It's never duplicated if the CI job already has it.
  • Some jobs use --platform amd/ascend (ROCm/Ascend runners) — resolve shows the platform; those won't install on an NVIDIA box.
  • Some jobs do extra setup inside the test step (e.g. cp .../maniskill_assets/assets into the repo, or export ROBOT_PLATFORM=ALOHA). The driver replays the whole CI test step, so those are included automatically — but they assume the asset dirs exist under /workspace/dataset/.
  • A duplicated config in list (e.g. d4rl_iql_mujoco,d4rl_iql_mujoco) just means the job runs that config twice with different flags (FSDP on/off). Normal.

Troubleshooting

  • No CI job for model=… env=… — that combo isn't in the workflow. Run list to see valid pairs; the model/env strings must match --model/--env in install.sh exactly (maniskill_libero, not libero).
  • AttributeError: 'MessageFactory' object has no attribute 'GetPrototype' in a test run — benign protobuf/TF-on-import noise from the workers, not a failure. Look for the Ray Placement(...) lines and the rollout progress bar to confirm real progress.
  • error: run this from inside the RLinf repo — cd to the repo root (the dir containing requirements/install.sh) before invoking the driver.
  • Test exits 0 immediately but nothing trained — you backgrounded it with &; the launcher returns 0 while training detaches. Run it in the foreground, or wait on the real train_embodied_agent.py pid.
  • install hangs at uv sync with ~0 CPU and no .uv_cache writes — almost always a dead http(s)_proxy env var on the box (e.g. a local 127.0.0.1:10809 that isn't forwarding), leaving idle ESTABLISHED :443 connections. unset http_proxy https_proxy HTTP_PROXY HTTPS_PROXY all_proxy ALL_PROXY before running the install. Not an install.sh bug.
  • setup_mirror fails: cannot overwrite multiple values ... insteadOf — prior interrupted runs left duplicate url.<mirror>.insteadOf entries in the global git config. Clear them with git config --global --unset-all url."https://gh-proxy.com/github.com/".insteadOf then retry. Also environmental, not an install.sh bug.

© RLinf, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in .agents/skills/test-install of RLinf/RLinf.

  • SKILL.md
  • driver.py

Open the folder on GitHubat commit c70606f

Compare with similar skills

Test Install next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Test Install compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Test Install this skillRLinf/RLinf5.4k—~2.3kAutomated safety check: PassApache-2.0
Web Application Testinganthropics/skills180k51 repos~966Automated safety check: PassApache-2.0
TDD WorkflowhellangleZ/burn-in-cceverywhere-ralph11211 repos~2.4kAutomated safety check: PassNone
Uloop Replay Inputkurotu/VRCQuestTools3733 repos~615Automated safety check: PassMIT
Ui4 Convert Testspayloadcms/payload45k—~3.5kAutomated safety check: PassMIT
E2Estackia/rtp2httpd2.2k—~517Automated safety check: PassGPL-2.0

Similar skills

  • Web Application Testing

    anthropics/skills

    Official

    Tests local web applications with Python Playwright scripts, checking frontend behavior, capturing screenshots and reading browser console logs.

    180k GitHub starsUsed in 51 repos~966 tokens
    Testing & QAAuto-check passed
  • TDD Workflow

    hellangleZ/burn-in-cceverywhere-ralph

    A skill your agent uses when writing new features, fixing bugs, or refactoring code.

    112 GitHub starsUsed in 11 repos~2.4k tokens
    Testing & QAAuto-check passed
  • Uloop Replay Input

    kurotu/VRCQuestTools

    Replay recorded PlayMode keyboard and mouse input. An agent skill from kurotu/VRCQuestTools.

    373 GitHub starsUsed in 3 repos~615 tokens
    Testing & QAAuto-check passed
  • Ui4 Convert Tests

    payloadcms/payload

    A skill your agent uses when UI changes are complete and e2e tests need updating.

    45k GitHub stars~3.5k tokensUpdated today
    Testing & QAAuto-check passed
  • E2E

    stackia/rtp2httpd

    Write, run, review, or debug rtp2httpd E2E tests and their harness in e2e/ and scripts/run-e2e.sh.

    2.2k GitHub stars~517 tokensUpdated 5 days ago
    Testing & QAAuto-check passed
  • Moav E2E

    MotherofallVPNs/MoaV

    Run and debug MoaV's end-to-end tests — real protocol connectivity (client-test.sh) and the moav CLI smoke test — against a LIVE server, via the self-hosted e2e workflow or a local test VPS.

    448 GitHub stars~1.9k tokensUpdated yesterday
    Testing & QAAuto-check: notes

More from RLinf/RLinf

All 9 skills in this repo
  • Adds example documentation for a new model or environment in RLinf (RST pages in the docs gallery for both English and Chinese).

    5.4k GitHub stars~1.9k tokensUpdated 5 days ago
    Auto-check passed
  • Adds a new publication page to the RLinf Sphinx docs (EN + ZH) and wires it into the Publications index/toctree.

    5.4k GitHub stars~1.1k tokensUpdated 5 days ago
    Auto-check passed
  • Create PR

    RLinf/RLinf

    Open a GitHub pull request for RLinf, or fix an existing one — checks the PR title against Conventional Commits, writes a precise description that follows .github/PULLREQUESTTEMPLATE.md, and lints…

    5.4k GitHub stars~2.2k tokensUpdated 5 days ago
    Auto-check passed
  • Docs Check

    RLinf/RLinf

    Cross-check RLinf documentation against code, natural explanation flow, and other docs, including English-Chinese parity.

    5.4k GitHub stars~2.1k tokensUpdated 5 days ago
    Auto-check passed
  • Install Check

    RLinf/RLinf

    Check, fix, or extend requirements/install.sh and its docker/Dockerfile coverage when adding a new embodied model or environment in RLinf, so the install logic reuses common utilities, keeps system…

    5.4k GitHub stars~2.3k tokensUpdated 5 days ago
    Auto-check passed
  • Adds install command in install script, Docker build stage in Dockerfile, and CI jobs for docker build and embodied e2e test when introducing a new model or environment in RLinf.

    5.4k GitHub stars~1.4k tokensUpdated 5 days ago
    Auto-check passed

Categories

Questions about Test Install

What does Test Install do?

Test that requirements/install.sh works for an embodied model/env by building its venv and running the matching CI e2e test. Test Install is an agent skill from RLinf/RLinf.sh works for an embodied model/env by building its venv and running the matching CI e2e test.

When should I use Test Install?

Test Install fits situations like: verify an install; check a new model/env installs cleanly; run the e2e test for a model/env; confirm a venv works.

How do I install Test Install in Claude Code?

Run `npx skills add RLinf/RLinf --skill test-install -a claude-code`. Or copy the skill folder (.agents/skills/test-install in RLinf/RLinf) into .claude/skills/test-install in your project. Claude Code loads it when a task matches its description.

How do I install Test Install in Codex?

Run `npx skills add RLinf/RLinf --skill test-install -a codex`. Or copy the skill folder (.agents/skills/test-install in RLinf/RLinf) into .agents/skills/test-install in your project. Codex loads it when a task matches its description.

Can I use Test Install in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add RLinf/RLinf --skill test-install -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/test-install, .gemini/skills/test-install, .github/skills/test-install and .opencode/skills/test-install in your project.

What does Test Install need to run?

Going by SKILL.md and its folder, Test Install needs Python for the scripts in its folder and the command-line tools its instructions call (python3, uv and git). Our summary lists: Python 3; Docker.

Does Test Install access the network?

SKILL.md names 1 domain. In commands or code: gh-proxy.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Test Install safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Test Install use?

Test Install is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Test Install use?

About 2.3k tokens (SKILL.md is roughly 9.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Test Install?

Skills that share tags, products or a category with Test Install: Web Application Testing (anthropics/skills, 180k stars), TDD Workflow (hellangleZ/burn-in-cceverywhere-ralph, 112 stars), Uloop Replay Input (kurotu/VRCQuestTools, 373 stars) and Ui4 Convert Tests (payloadcms/payload, 45k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Test Install?

RLinf (a GitHub organization) maintains it in RLinf/RLinf, which has 5,447 GitHub stars. The repository holds 9 skills in this directory. The repository was last updated on October 2, 2026.

Source: RLinf/RLinf on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.