Arize Evaluator
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
Front-door for COMPASS — training, evaluation, SAGE scene workflows (search / USD conversion / scene registration), and OSMO cloud submission.
$ npx skills add NVlabs/COMPASS --skill compass -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install NVlabs/COMPASS compass --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/NVlabs/COMPASS.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/compass .claude/skills/compass && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "compass" agent skill from https://github.com/NVlabs/COMPASS/tree/main/.claude/skills/compass into .claude/skills/compass/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compass", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/NVlabs/COMPASS/tree/main/.claude/skills/compassType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add NVlabs/COMPASS --skill compass -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install NVlabs/COMPASS compass --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVlabs/COMPASS.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/compass .agents/skills/compass && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "compass" agent skill from https://github.com/NVlabs/COMPASS/tree/main/.claude/skills/compass into .agents/skills/compass/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compass", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVlabs/COMPASS --skill compass -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install NVlabs/COMPASS compass --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVlabs/COMPASS.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/compass .cursor/skills/compass && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "compass" agent skill from https://github.com/NVlabs/COMPASS/tree/main/.claude/skills/compass into .cursor/skills/compass/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compass", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/NVlabs/COMPASS.git --path .claude/skills/compass--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add NVlabs/COMPASS --skill compass -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install NVlabs/COMPASS compass --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVlabs/COMPASS.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/compass .gemini/skills/compass && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "compass" agent skill from https://github.com/NVlabs/COMPASS/tree/main/.claude/skills/compass into .gemini/skills/compass/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compass", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install NVlabs/COMPASS compassInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add NVlabs/COMPASS --skill compass -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/NVlabs/COMPASS.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/compass .github/skills/compass && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "compass" agent skill from https://github.com/NVlabs/COMPASS/tree/main/.claude/skills/compass into .github/skills/compass/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compass", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add NVlabs/COMPASS --skill compass -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install NVlabs/COMPASS compass --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/NVlabs/COMPASS.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/compass .opencode/skills/compass && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "compass" agent skill from https://github.com/NVlabs/COMPASS/tree/main/.claude/skills/compass into .opencode/skills/compass/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "compass", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
compassFront-door for COMPASS — training, evaluation, SAGE scene workflows (search / USD conversion / scene registration), and OSMO cloud submission.
Compass is an agent skill from NVlabs/COMPASS. Front-door for COMPASS — training, evaluation, SAGE scene workflows (search / USD conversion / scene registration), and OSMO cloud submission. Use whenever the user mentions training a policy, adding a SAGE scene, evaluating a checkpoint, or running COMPASS in general. For debug / onboarding-a-new-robot, see the specialty siblings: compass-doctor, compass-newembodiment.
Its SKILL.md is about 5.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files, including scripts and reference files (for example `references/setup-sage-local.md`, `scripts/sage10k_search.py` and `scripts/sage10k_to_usd.py`).
The repository describes itself as: Cross-embOdiment Mobility Policy via ResiduAl RL and Skill Synthesis. The licence is Apache-2.0.
8 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 8060f7c. It shows what the files ask for, not the result of running them.
Pre-approves these tools, so the agent can use them without asking each time:
BashReadEditWriteGrepGlobAskUserQuestionWebFetchAgentFrom allowed-tools in the SKILL.md frontmatter.
Ships 2 files in scripts/ (Python), which the agent can run.
Shell commands in SKILL.md call:
pythondockerhuggingface-clipython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.coAlso links to:
nvlabs.github.ioFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
SCENE_KEYHF_TOKENWANDB_API_KEYFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Compass loads about 5.4k tokens when it runs, and up to ~5.8k if it reads all its reference files. Until then it costs about 95 tokens; SKILL.md has 2,078 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check noted patterns worth knowing about, such as sudo or a known installer.
allowed-tools: Bash, Read, Edit, Write, Grep, Glob, AskUserQuestion, WebFetch, AgentAutomated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from NVlabs/COMPASS at commit 8060f7c, republished under its Apache-2.0 licence (© NVlabs). 2,078 words, ~5,407 tokens.
.claude/skills/compass/SKILL.md (or your agent's skills folder). This skill also uses 3 other files; get the full folder from GitHub.You orchestrate the COMPASS robot navigation training pipeline — from scene search to trained policy. Adapt to context: interactive and explanatory during setup, autonomous and efficient for repeated operations like training runs.
COMPASS trains cross-embodiment navigation policies using residual RL on top of X-Mobility (a pretrained VLA). SAGE-10k provides 10,000 pre-generated indoor scenes across 50 room types.
This skill bundles two helper scripts in its own directory:
scripts/sage10k_search.py — search SAGE-10k dataset by text queryscripts/sage10k_to_usd.py — convert SAGE-10k scenes to USD (uses only Isaac Sim native APIs, no extra deps)Use the full path when invoking them:
<SKILL_BASE_DIR>/scripts/sage10k_search.py
<SKILL_BASE_DIR>/scripts/sage10k_to_usd.py| Workflow | Trigger |
|---|---|
| Search SAGE-10k | User wants a scene from the SAGE-10k dataset (default for new scenes) |
| Setup COMPASS | User wants to install COMPASS deps, or deps are missing when needed |
| Setup SAGE (local) | User wants full local SAGE installation for custom generation (rare; see references/setup-sage-local.md) |
| Register Scene | User has a USD file to add to COMPASS (auto-triggered after conversion) |
| Train | User wants to run residual RL training on a scene |
| Evaluate | User wants to evaluate a trained checkpoint |
| Full Pipeline | User gives a scene description — chain: search SAGE-10k → download → convert USD → verify in viewer → register → preview train (num_envs=1, GUI) → full train |
If the user gives a scene description like "a cluttered warehouse" or "bedroom", treat it as Full Pipeline using SAGE-10k search. Use local SAGE only when the user explicitly asks for it.
Before any of Train, Evaluate, or Full Pipeline, run the Prerequisites Check first. The cost of a stale check is a confusing CUDA error 30 minutes into a training run.
The Full Pipeline has two visual checkpoints that you MUST NOT skip, even when the user has told you to work without clarifying questions:
--viz kit preview training (see Train → Step 1: Preview training) — launch with num_envs=1 --viz kit so the Kit window shows the robot in the scene, and wait for the user's confirmation before launching the full-scale headless run."Work without clarifying questions" applies to ambiguities about what to do (scene choice, embodiment, num_envs). It does not override these two human-eye verification handoffs. SAGE conversions and Isaac Lab spawn settings can produce silently-broken scenes (rooms missing 3 of 4 walls, robot spawned outside the room, scene upside-down from wrong quaternion) that no programmatic check catches — only a human watching the viewer does. Skipping either step usually costs hours of wasted GPU on a broken scene.
Some user intents are better handled by sibling skills. When the user's task fits one of these, say so to the user and recommend the matching specialty — don't try to handle it inside compass:
| User intent | Specialty skill |
|---|---|
| Diagnose why training won't start, "what's wrong", quick health check | /compass-doctor |
| Add a new robot platform (cfg files, EmbodimentEnvCfgMap registration) | /compass-newembodiment |
Tell the user the matching specialty and suggest they invoke it (or rephrase their ask so the auto-router picks it up). Don't programmatically invoke the sibling via the Skill tool — that adds latency and removes the user's ability to redirect.
Run this before any operation. The skill assumes the user has run source ./docker/activate (the docker-as-venv shim from the docker/ subdir). Run all checks in parallel where possible. Report what's ready and what's missing.
# Container running?
./docker/run.sh status
# Activated shell? `deactivate` is a shell function defined by ./docker/activate
command -v deactivate >/dev/null && echo "shell: activated" || echo "shell: NOT activated — run: source ./docker/activate"
# GPU (use dangerouslyDisableSandbox: true)
nvidia-smi --query-gpu=name,memory.total,memory.used --format=csv,noheader
# Assets present
test -f ./assets/x_mobility.ckpt && echo "x_mobility ckpt: OK ($(du -h ./assets/x_mobility.ckpt | cut -f1))"
ls ./assets/usd/ 2>/dev/null | head -5If anything fails, route to Setup COMPASS. If the user reports a vague "training won't start" issue, that's compass-doctor's job — recommend it.
These rules apply to every command. Each one exists for a specific reason; the explanation matters more than the rule itself, because edge cases need judgment.
Activated-shell rule. Every Python invocation runs inside a shell where the user has already done source ./docker/activate. Inside that shell, python, pip, tensorboard, etc. are shims that route into the COMPASS container automatically. The right command is just python run.py … — no Isaac Lab launcher prefix, no conda env wrapper. If command -v deactivate returns nothing, the shell isn't activated — pause and ask the user to source the activate script before continuing.
GPU access needs dangerouslyDisableSandbox: true. The Claude Code sandbox blocks NVIDIA driver access — without this flag, nvidia-smi fails and Isaac Sim can't find CUDA. This is a Claude Code concern, not a container concern, so it applies even though the GPU work happens inside the container.
Background execution for long jobs. Training and evaluation can run for hours. Use run_in_background: true so the user isn't blocked. To check progress, look at log files instead of stdout (the shim forwards stdout but it's easy to miss when the run is backgrounded):
# Find latest training log
ls -t /tmp/isaaclab/logs/ | head -1
# Or check the Kit log for errors
find ~/.local/share/ov/pkg/isaac-sim-* -name "kit_*.log" 2>/dev/null | head -1 | xargs tail -100Killing processes. kill -9 <PID> of the python process. The shim forwards SIGTERM cleanly into the container, no orphan-conda-wrapper issues to worry about:
ps aux | grep "run.py.*<scene>" | grep -v grep | awk '{print $2}' | xargs kill -9The new dev environment uses Docker as a venv: build the image once, download assets once, then activate. Three commands:
export HF_TOKEN=hf_xxx # https://huggingface.co/settings/tokens
./docker/run.sh build # ~10 min cold build of compass-rl image
./docker/run.sh assets # ~5 min: downloads compass_usds + x_mobility ckpt
source ./docker/activate # prompt becomes (compass-rl)After this, python run.py … (and pip, tensorboard, etc.) Just Work — they route to the container via docker exec. No conda env, no manual pip-install of requirements, no Isaac Lab clone. Assets land at ./assets/usd/ (built-in scenes) and ./assets/x_mobility.ckpt (base policy), bind-mounted into the container at the same paths under /workspace/COMPASS/.
Full reference: docs/handbook/installation/docker.md. If the user wants the bare-metal install (no Docker), point them at docs/handbook/installation/bare-metal.md — supported but slower to set up and not the recommended default.
The SAGE-10k dataset (nvidia/SAGE-10k) has 10,000 pre-generated indoor scenes across 50 room types. No SAGE installation needed.
python <SKILL_BASE_DIR>/scripts/sage10k_search.py "<USER_PROMPT>" --top 5 --sample-size 50First run builds a scene index from the HF API (cached afterward). Results show room type, style, object count, and description.
Show top results. Let user pick one.
huggingface-cli download nvidia/SAGE-10k --repo-type dataset \
--include "scenes/<zip_filename>" --local-dir ./sage_10k_cache
mkdir -p ./sage_10k_scenes/<scene_name>
unzip ./sage_10k_cache/scenes/<zip_filename> -d ./sage_10k_scenes/<scene_name>/The bundled converter only needs Isaac Sim native APIs, no extra deps:
# dangerouslyDisableSandbox: true (Isaac Sim needs GPU)
python <SKILL_BASE_DIR>/scripts/sage10k_to_usd.py \
./sage_10k_scenes/<scene_name>/<layout_json> \
./compass/rl_env/exts/mobility_es/mobility_es/usd/<scene_name>/<scene_name>.usdThe output USD lives under the mobility_es extension directory because that's where environments.py expects to find it (the registered USD_PATHS dict resolves ../usd/<scene>/... relative to the config/ subdir). Built-in COMPASS scenes from ./assets/usd/ are referenced separately; user-added SAGE scenes go here.
The script:
pxr accessThis step is not optional and cannot be skipped, including under "no clarifying questions" mode — see Mandatory visual-verification gates.
Launch the Isaac Sim viewer on the converted USD so the user can verify the scene looks correct (geometry, collisions, textures). Use dangerouslyDisableSandbox: true and run_in_background: true:
python -c "
from isaacsim import SimulationApp
app = SimulationApp({'headless': False, 'width': 1280, 'height': 720})
import omni
omni.usd.get_context().open_stage('./compass/rl_env/exts/mobility_es/mobility_es/usd/<scene_name>/<scene_name>.usd')
while app.is_running():
app.update()
app.close()
" &Then:
:1 and to inspect it.If they describe an issue, diagnose it (re-convert, adjust converter, etc.) and re-open the viewer for another round of verification. Repeat until they confirm.
Only after the user has explicitly confirmed the USD in Step 5, proceed to Register Scene → Train.
Edit two files. Read each first to find the right insertion point — line numbers drift over time.
compass/rl_env/exts/mobility_es/mobility_es/config/environments.pyAdd to USD_PATHS dict:
'<SceneName>':
os.path.join(os.path.dirname(__file__), "../usd/<scene_dir>/<scene_file>.usd"),Add scene config at end of file:
<scene_var> = EnvSceneAssetCfg(
prim_path="{ENV_REGEX_NS}/<SceneName>",
init_state=AssetBaseCfg.InitialStateCfg(
pos=(0, 0, 0.01),
# Isaac Lab 3.0 quaternion is (x, y, z, w) — w LAST. Identity is
# (0,0,0,1). Do NOT use (1,0,0,0): that's a 180° flip about X and
# turns the scene upside-down. Existing scenes in this file use
# (0,0,0,1) for the same reason.
rot=(0.0, 0.0, 0.0, 1.0),
),
spawn=sim_utils.UsdFileCfg(
usd_path=USD_PATHS['<SceneName>'],
scale=(1.0, 1.0, 1.0),
rigid_props=sim_utils.RigidBodyPropertiesCfg(
disable_gravity=None,
solver_position_iteration_count=4,
solver_velocity_iteration_count=1,
),
),
# IMPORTANT: SAGE-10k rooms are NOT centered on origin — their walls span
# absolute layout coordinates, e.g. x:(0..3.7), y:(0..3). pose_sample_range
# is in env-local coords, so set it to match the actual wall bounds with
# a ~0.5m safety margin inside the walls. Inspect the converted USD's
# bbox before picking values; "symmetric ±N" only works if the room
# happens to be centered at origin.
pose_sample_range={"x": (0.5, 3.2), "y": (0.5, 2.5), "yaw": (-3.14, 3.14)},
env_spacing=20,
)For SAGE-10k rooms (typically 4–6m wall-to-wall), env_spacing=20 provides plenty of clearance between parallel envs. Always compute pose_sample_range from the actual wall extent rather than assuming the room is centered.
Quick way to dump the room's world bbox (run inside the activated shell):
from pxr import Usd, UsdGeom
stage = Usd.Stage.Open("<path/to/scene>.usd")
bbox = UsdGeom.BBoxCache(Usd.TimeCode.Default(),
includedPurposes=[UsdGeom.Tokens.default_]
).ComputeWorldBound(stage.GetPseudoRoot()
).ComputeAlignedRange()
print(bbox.GetMin(), bbox.GetMax())run.pyAdd to EnvSceneAssetCfgMap (around line 118–128):
'<scene_var>': environments.<scene_var>,This step is not optional and cannot be skipped, including under "no clarifying questions" mode — see Mandatory visual-verification gates.
Launch with --num_envs 1 and --viz kit so the Isaac Sim Kit window opens and the user can verify the robot spawns correctly in the scene. Use run_in_background: true:
python run.py \
-c configs/train_config.gin \
--enable_cameras \
--viz kit \
-o <OUTPUT_DIR>_preview \
-b ./assets/x_mobility.ckpt \
-n <WANDB_PROJECT> \
-r <RUN_NAME>_preview \
--logger tensorboard \
--video \
--embodiment <EMBODIMENT> \
--environment <SCENE_KEY> \
--num_envs 1Notes on flags:
--viz kit opens the Kit viewer. Isaac Lab 3.0 deprecated --headless; headless is the new default, and you opt in to a visualizer with --viz {kit,newton,rerun,viser} (or --viz none to be explicit). Do not pass --headless.-n/--wandb-project-name is required by run.py even when --logger tensorboard — any string works.--viz kit, expect ~30% slower throughput (renderer overhead) and ~9 GB extra GPU than headless.Tell the user the preview is running with the GUI open on display :1. Ask them to verify:
Stop and wait for the user's explicit confirmation before proceeding to Step 2. Do not auto-progress to full-scale training. If the user describes an issue, kill the preview, fix it (USD re-convert, pose_sample_range adjustment, embodiment swap, etc.), and re-launch the preview for another round.
After preview confirmation, launch the full run with run_in_background: true:
python run.py \
-c configs/train_config.gin \
--enable_cameras \
-o <OUTPUT_DIR> \
-b ./assets/x_mobility.ckpt \
-n <WANDB_PROJECT> \
-r <RUN_NAME> \
--logger tensorboard \
--video \
--embodiment <EMBODIMENT> \
--environment <SCENE_KEY> \
--num_envs <NUM_ENVS>Headless is the default in Isaac Lab 3.0 — omit --viz entirely (or use --viz none to be explicit). Do not pass --headless; it's deprecated. If you want a viewer for a brief look, add --viz kit and drop --num_envs to 1 (renderer + many envs will OOM, especially on SAGE-10k scenes).
The Python launcher in osmo/run_osmo.py handles build+push+submit and is the recommended entry point.
run_osmo.py is host-side — it shells out to docker build, docker push, and osmo workflow submit, none of which exist inside the COMPASS runtime container. The python shim from source ./docker/activate recognizes this via a # COMPASS_HOST_SIDE: true marker at the top of the launcher and auto-routes it to host Python, so plain python osmo/run_osmo.py … works from the activated shell. If the user sourced the activate script before the marker-aware shim shipped, have them deactivate && source ./docker/activate to refresh, or fall back to /usr/bin/python3 osmo/run_osmo.py ….
export WANDB_API_KEY=<key>
export HF_TOKEN=<token>
export COMPASS_OSMO_REGISTRY=nvcr.io/<org>/<team>
python osmo/run_osmo.py train \
--experiment-name <name> \
--wandb-project <project>The X-Mobility base checkpoint is now downloaded inside the workflow from huggingface.co/nvidia/X-Mobility, so no --base-policy-ckpt flag is needed. For direct osmo workflow submit invocation, see the OSMO cloud submission handbook page. Workflow YAMLs live in osmo/workflows/.
| Parameter | Default | Source |
|---|---|---|
| num_iterations | 1000 | train_config.gin |
| num_envs | 32 | train_config.gin |
| num_steps_per_iteration | 256 | shared.gin |
| seed | 20 | train_config.gin |
| embodiment | g1 | shared.gin |
carter, h1, spot, g1, digit
warehouse_single_rack, galileo_lab, simple_office, combined_single_rack, combined_multi_rack, random_envs, hospital, warehouse_multi_rack
<OUTPUT_DIR>/model_*.pt<OUTPUT_DIR>/videos/<OUTPUT_DIR>/ or W&BCKPT=$(ls <TRAIN_OUTPUT_DIR>/model_*.pt | sort -V | tail -n 1)
python run.py \
-c configs/eval_config.gin \
--enable_cameras \
-o <OUTPUT_DIR> \
-b ./assets/x_mobility.ckpt \
-p $CKPT \
-n <WANDB_PROJECT> \
-r eval_<RUN_NAME> \
--logger tensorboard \
--video \
--video_interval 1 \
--embodiment <EMBODIMENT> \
--environment <SCENE_KEY>Headless by default (Isaac Lab 3.0). Add --viz kit to open the viewer for visual inspection of episodes. Do not pass --headless — deprecated.
To inspect a USD scene before training (check collision meshes, layout, etc.):
python -c "
from isaacsim import SimulationApp
app = SimulationApp({'headless': False, 'width': 1280, 'height': 720})
import omni
omni.usd.get_context().open_stage('<USD_PATH>')
while app.is_running():
app.update()
app.close()
" &Use dangerouslyDisableSandbox: true. Suggest the user enable Show > Physics > Colliders to verify collision meshes.
Path convention: <USD_PATH> must be a path the container can see — pass either a path relative to the repo root (e.g. ./compass/rl_env/.../scene.usd) or the container-side absolute path (/workspace/COMPASS/...). Do not pass the host absolute path (/home/<user>/Projects/COMPASS/...) — the container has the repo bind-mounted at /workspace/COMPASS and will fail to open a host path with "Failed to get crate info from file".
Most users do not need this. SAGE-10k search (above) covers the typical case without any SAGE install. Only use a local SAGE install when the user explicitly wants to generate fully custom scenes (novel layouts not in SAGE-10k).
For installation steps, dependencies, and trade-offs, see references/setup-sage-local.md — that file is loaded only when a local SAGE install is actually needed, to keep the main flow lean.
| File | Purpose |
|---|---|
docker/run.sh | Build / assets / up / down / exec / shell / status |
docker/activate | Sourceable activate script (sets up python/pip shims) |
./assets/x_mobility.ckpt | X-Mobility base policy (downloaded by ./docker/run.sh assets) |
./assets/usd/ | Built-in scene USDs (downloaded by ./docker/run.sh assets) |
run.py | Main entry point + EnvSceneAssetCfgMap (around lines 118–128) |
compass/rl_env/exts/mobility_es/mobility_es/config/environments.py | Scene definitions, USD_PATHS, EnvSceneAssetCfg class |
compass/rl_env/exts/mobility_es/mobility_es/usd/ | User-added scene USDs (e.g., SAGE-10k conversions) |
configs/train_config.gin | Training parameters |
configs/eval_config.gin | Evaluation parameters |
configs/shared.gin | Shared config (embodiment, environment, steps) |
osmo/workflows/rl_es_train_workflow.yaml | OSMO training workflow template |
osmo/run_osmo.py | Python launcher for OSMO submission (train / eval / record / distill) |
<entity>/<project>/<artifact>:<version> — e.g. <your-entity>/<your-project>/x_mobility:v1
© NVlabs, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 3 other files (scripts, references) in .claude/skills/compass of NVlabs/COMPASS.
Open the folder on GitHubat commit 8060f7c
Compass next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Compass this skillNVlabs/COMPASS | 145 | — | ~5.4k | Automated safety check: Notes | Apache-2.0 | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 2 repos | ~8.1k | Automated safety check: Notes | MIT | |
| LLM Evaluationdavila7/claude-code-templates | 32k | 13 repos | ~3.5k | Automated safety check: Pass | MIT | |
| Train Poseruvnet/RuView | 97k | — | ~504 | Automated safety check: Pass | MIT | |
| Agent Evaluationsickn33/agentic-awesome-skills | 47k | 1 repos | ~2k | Automated safety check: Pass | MIT | |
| Ray Train Distributed TrainingOrchestra-Research/AI-Research-SKILLs | 13k | 3 repos | ~2.7k | Automated safety check: Pass | MIT |
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
ruvnet/RuView
Train/evaluate WiFi pose models honestly — camera-supervised (MediaPipe + CSI) and camera-free (WiFlow), always checked against the mean-pose baseline before any PCK is quoted.
sickn33/agentic-awesome-skills
Evaluate agent behavior with versioned cases and explicit verifiers.
Orchestra-Research/AI-Research-SKILLs
Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.
Arize-ai/phoenix
Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.
NVlabs/COMPASS
Diagnose why COMPASS isn't working: container, GPU, activated shell, assets, Isaac Sim init, checkpoint validity.
NVlabs/COMPASS
Onboard a new robot platform to COMPASS: generate robot ArticulationCfg + per-embodiment envcfg, register in EmbodimentEnvCfgMap, smoke-test the spawn with --numenvs 1.
Front-door for COMPASS — training, evaluation, SAGE scene workflows (search / USD conversion / scene registration), and OSMO cloud submission. Compass is an agent skill from NVlabs/COMPASS. Front-door for COMPASS — training, evaluation, SAGE scene workflows (search / USD conversion / scene registration), and OSMO cloud submission.
Compass fits situations like: the user mentions training a policy; adding a SAGE scene; evaluating a checkpoint; running COMPASS in general.
Run `npx skills add NVlabs/COMPASS --skill compass -a claude-code`. Or copy the skill folder (.claude/skills/compass in NVlabs/COMPASS) into .claude/skills/compass in your project. Claude Code loads it when a task matches its description.
Run `npx skills add NVlabs/COMPASS --skill compass -a codex`. Or copy the skill folder (.claude/skills/compass in NVlabs/COMPASS) into .agents/skills/compass in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVlabs/COMPASS --skill compass -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/compass, .gemini/skills/compass, .github/skills/compass and .opencode/skills/compass in your project.
Going by SKILL.md and its folder, Compass needs Python for the scripts in its folder, the command-line tools its instructions call (python, docker, huggingface-cli and python3) and credentials named SCENE_KEY, HF_TOKEN and WANDB_API_KEY. Our summary lists: Python 3; Docker. Its frontmatter pre-approves these tools: Bash, Read, Edit, Write, Grep, Glob, AskUserQuestion, WebFetch, Agent.
SKILL.md names 2 domains. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. As links in the text: nvlabs.github.io. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found notes only (pre-approves every shell command (allowed-tools: bash)), nothing it rates as a warning. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Compass is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 5.4k tokens (SKILL.md is roughly 22k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 435 tokens, read only when the agent opens those files.
Skills that share tags, products or a category with Compass: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Train Pose (ruvnet/RuView, 97k stars) and Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
NVlabs (a GitHub organization) maintains it in NVlabs/COMPASS, which has 145 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on October 7, 2026.
Source: NVlabs/COMPASS on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.