Policy Monitor
anthropics/claude-for-legal
Keep the AI policy current with practice — weekly sweep of saved AIAs, triage results, and vendor reviews to find policy drift, or direct query for a proposed new AI practice.
Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.
$ npx skills add InternRobotics/EBench --skill ebench-evaluate -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install InternRobotics/EBench ebench-evaluate --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ebench-evaluate .claude/skills/ebench-evaluate && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "ebench-evaluate" agent skill from https://github.com/InternRobotics/EBench/tree/main/skills/ebench-evaluate into .claude/skills/ebench-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ebench-evaluate", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/InternRobotics/EBench/tree/main/skills/ebench-evaluateType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add InternRobotics/EBench --skill ebench-evaluate -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install InternRobotics/EBench ebench-evaluate --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .agents/skills && cp -r skills-src/skills/ebench-evaluate .agents/skills/ebench-evaluate && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "ebench-evaluate" agent skill from https://github.com/InternRobotics/EBench/tree/main/skills/ebench-evaluate into .agents/skills/ebench-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ebench-evaluate", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add InternRobotics/EBench --skill ebench-evaluate -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install InternRobotics/EBench ebench-evaluate --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/skills/ebench-evaluate .cursor/skills/ebench-evaluate && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "ebench-evaluate" agent skill from https://github.com/InternRobotics/EBench/tree/main/skills/ebench-evaluate into .cursor/skills/ebench-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ebench-evaluate", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/InternRobotics/EBench.git --path skills/ebench-evaluate--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add InternRobotics/EBench --skill ebench-evaluate -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install InternRobotics/EBench ebench-evaluate --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/skills/ebench-evaluate .gemini/skills/ebench-evaluate && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "ebench-evaluate" agent skill from https://github.com/InternRobotics/EBench/tree/main/skills/ebench-evaluate into .gemini/skills/ebench-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ebench-evaluate", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install InternRobotics/EBench ebench-evaluateInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add InternRobotics/EBench --skill ebench-evaluate -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .github/skills && cp -r skills-src/skills/ebench-evaluate .github/skills/ebench-evaluate && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "ebench-evaluate" agent skill from https://github.com/InternRobotics/EBench/tree/main/skills/ebench-evaluate into .github/skills/ebench-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ebench-evaluate", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add InternRobotics/EBench --skill ebench-evaluate -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install InternRobotics/EBench ebench-evaluate --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/skills/ebench-evaluate .opencode/skills/ebench-evaluate && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "ebench-evaluate" agent skill from https://github.com/InternRobotics/EBench/tree/main/skills/ebench-evaluate into .opencode/skills/ebench-evaluate/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "ebench-evaluate", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
ebench-evaluateRun and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.
Ebench Evaluate is an agent skill from InternRobotics/EBench. Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.
Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
The repository describes itself as: Elemental Diagnosis of Generalist Mobile Manipulation Policies. The licence is MIT.
4 steps, taken from the first numbered list in SKILL.md.
Read from SKILL.md and the folder at commit 355fe56. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
jqbashFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Ebench Evaluate loads about 1.9k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 931 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from InternRobotics/EBench at commit 355fe56, republished under its MIT licence (© InternRobotics). 931 words, ~1,925 tokens.
.claude/skills/ebench-evaluate/SKILL.md (or your agent's skills folder).Work from the EBench root. Read the selected baseline entry point and the relevant CLI implementation under third_party/genmanip-client/src/genmanip_client/. Verify installed gmp ... --help before relying on flags from a different revision.
Resolve model/checkpoint, track, split, server mode, GPU/worker budget, and output directory from the request. Use val_train / val_unseen for tuning. Run held-out test when requested for final evaluation; do not silently substitute a split or tune on held-out results.
Record a small manifest beside the run logs: EBench and submodule commits, local code changes, checkpoint revision/path, config and normalization source, track/split/task selection, run/task ID, worker-to-GPU mapping, action/replan horizon, start time, sanitized command, and result locations. Mark unavailable fields as unknown. Exclude tokens and signed credentials.
gmp submit "$CONFIG_PATH" --run_id "$RUN_ID" --host "$SERVER_HOST" --port "$SERVER_PORT". The pinned submit CLI takes host/port, not the online platform's --base_url. Submission schedules jobs; it does not load the user's policy.extensions/online_cli.py; gmp online submit --base_url "$PLATFORM_URL" --token "$TOKEN" --model_name "$MODEL_NAME" --benchmark_set EBench --timeout 600 --print_endpoint creates a task and waits for readiness. Choose metadata, visibility and wait budget appropriate to the user's request; 600 seconds is an example budget, not a service guarantee. Capture returned task_id and endpoint; use that task ID as the client run ID.gmp online ready --base_url "$PLATFORM_URL" --token "$TOKEN" --task_id "$RUN_ID"; use its returned evaluation endpoint. A wait timeout does not prove creation failed. Inspect existing task state before retrying creation; do not create duplicate tasks while resources are pending.The platform URL, returned evaluation endpoint, and an OpenPI model server address serve different purposes. Do not interchange them. Preserve credentials through local environment/configuration, omit them from reports, and avoid shell tracing of authenticated commands.
For a new online evaluation, the agent should complete this sequence rather than require the user to obtain an endpoint manually. First verify the local model environment/checkpoint and obtain the platform URL and locally configured token. A new task does not need a pre-existing evaluation URL or task ID.
Submit and wait in the queue. gmp online submit creates the task and polls until evaluation resources are ready. Run it as a monitored process; while pending, report that the task is waiting, not evaluating. This Bash example requires jq and uses a configurable wait budget:
set -euo pipefail
: "${PLATFORM_URL:?Set the online platform URL}"
: "${TOKEN:?Set the API token locally}"
: "${MODEL_NAME:?Set the model name}"
READY_JSON=$(gmp online submit \
--base_url "$PLATFORM_URL" \
--token "$TOKEN" \
--model_name "$MODEL_NAME" \
--model_type VLA \
--benchmark_set EBench \
--timeout "${QUEUE_TIMEOUT_SECONDS:-600}" \
--print_endpoint)
# Parse only a successful ready response; never launch with empty values.
EVAL_URL=$(printf '%s' "$READY_JSON" | jq -er '.endpoint | strings | select(length > 0)')
RUN_ID=$(printf '%s' "$READY_JSON" | jq -er '.task_id | strings | select(length > 0)')
export EVAL_URL RUN_ID--print_endpoint returns a JSON object containing both endpoint and task_id, not a plain URL. If submission fails, times out, or either field is missing, stop before launching the client. A timeout may leave a task queued: recover its ID from available logs/platform state and query gmp online ready instead of submitting again. If its ID cannot be determined, report that uncertainty rather than create a duplicate.
Save the assignment. Record the returned task ID and endpoint in the local run metadata, excluding credentials. Use the returned task ID unchanged as RUN_ID; do not substitute a friendly experiment name. Keep any credential-bearing endpoint out of shared reports.
Start actual model inference against the assigned endpoint. Use the baseline commands below or the custom adapter. Pass EVAL_URL as the evaluation server address and RUN_ID as the run ID. Do not run a second gmp submit against the online endpoint: the online task already schedules the evaluation. For OpenPI, ensure its separate local model server is ready before launching the eval client.
Monitor until evaluation completes. Use the assigned endpoint/task ID for status and preserve the resulting logs and episode artifacts. Queue readiness only means resources are available; it does not mean the model has been evaluated.
An existing ready task starts at step 2; an existing queued task uses gmp online ready until ready within the chosen wait budget. Once both fields are valid, continue to evaluation within the user's request without asking them to copy the values back manually.
gmp eval supplies fake actions. Use it only for an explicitly scoped connectivity smoke test, never as evidence of a checkpoint's performance.
For X-VLA, the existing wrapper accepts environment variables (here EVAL_URL is the evaluation endpoint):
MODEL_PATH="$CHECKPOINT" BASE_URL="$EVAL_URL" RUN_ID="$RUN_ID" \
TOKEN="$TOKEN" WORKER_IDS=0 GPU_IDS=0 LOG_DIR="$RUN_LOG_DIR" \
bash scripts/run_xvla_eval.shFor OpenPI, use the baseline README plus the actual serve_policy.py and pi_eval_client_online.py arguments; resolve the launch template's placeholders first. The current client constructs its policy adapter for worker_ids[0], so launch one client process per worker rather than assuming one process drives all listed workers.
For InternVLA-A1, run inference.py from its baseline directory with its supported --ckpt_path, --url, --run_id, --token, and --worker_ids arguments. Check the implementation before relying on a wrapper mentioned only in documentation.
Across hosts, share the run ID and assign disjoint worker IDs. A GPU index is not a worker ID. Start with a small validation smoke run for a newly integrated policy before expanding to the requested budget; do not submit an extra held-out smoke run by default.
Use gmp status --url "$EVAL_URL" --run_id "$RUN_ID" --token "$TOKEN" with worker logs. Distinguish waiting for resources, model loading, active steps, completed episodes, and transport failures. Bound polling/recovery to the user's time budget. Do not overwrite runs, clean results, or restart unrelated workers as routine recovery.
At completion, verify server status and expected versus saved episode coverage, not just process exit code. Capture failed/missing episodes and interruptions. Locate actual client outputs (default client_results, overridable by GENMANIP_RESULT_DIR) and server outputs separately. Report completion or partial completion, run ID, manifest/log/result paths, and any blocker. A request to run evaluation does not by itself request separate leaderboard publication.
© InternRobotics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in skills/ebench-evaluate of InternRobotics/EBench.
Open the folder on GitHubat commit 355fe56
Ebench Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Ebench Evaluate this skillInternRobotics/EBench | 145 | — | ~1.9k | Automated safety check: Pass | MIT | |
| Policy Monitoranthropics/claude-for-legal | 9.6k | 3 repos | ~3.7k | Automated safety check: Pass | Apache-2.0 | |
| Cosmos Policy EvaluationOrchestra-Research/AI-Research-SKILLs | 13k | — | ~3.7k | Automated safety check: Pass | MIT | |
| Policy Monitoranthropics/claude-for-legal | 9.6k | 2 repos | ~4.5k | Automated safety check: Pass | Apache-2.0 | |
| Implementing Policy As Code With Open Policy Agentmukul975/Anthropic-Cybersecurity-Skills | 34k | — | ~2.6k | Automated safety check: Notes | Apache-2.0 | |
| Arize Evaluatorgithub/awesome-copilot | 40k | 2 repos | ~8.1k | Automated safety check: Notes | MIT |
anthropics/claude-for-legal
Keep the AI policy current with practice — weekly sweep of saved AIAs, triage results, and vendor reviews to find policy drift, or direct query for a proposed new AI practice.
Orchestra-Research/AI-Research-SKILLs
Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.
anthropics/claude-for-legal
Keep the privacy policy current with practice. An agent skill from anthropics/claude-for-legal.
mukul975/Anthropic-Cybersecurity-Skills
Implements policy-as-code enforcement with Open Policy Agent (OPA) and Gatekeeper for Kubernetes and CI/CD pipelines, covering writing Rego policies, deploying OPA Gatekeeper as a Kubernetes…
github/awesome-copilot
Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…
davila7/claude-code-templates
Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.
InternRobotics/EBench
Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.
InternRobotics/EBench
Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.
InternRobotics/EBench
Prepare or check an EBench evaluation environment for OpenPI, X-VLA, InternVLA-A1, or a custom policy.
InternRobotics/EBench
Diagnose EBench evaluation failures, stalled workers, transport errors, invalid actions, and unexpectedly low scores using logs and episode artifacts.
Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records. Ebench Evaluate is an agent skill from InternRobotics/EBench. Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.
Run `npx skills add InternRobotics/EBench --skill ebench-evaluate -a claude-code`. Or copy the skill folder (skills/ebench-evaluate in InternRobotics/EBench) into .claude/skills/ebench-evaluate in your project. Claude Code loads it when a task matches its description.
Run `npx skills add InternRobotics/EBench --skill ebench-evaluate -a codex`. Or copy the skill folder (skills/ebench-evaluate in InternRobotics/EBench) into .agents/skills/ebench-evaluate in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add InternRobotics/EBench --skill ebench-evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ebench-evaluate, .gemini/skills/ebench-evaluate, .github/skills/ebench-evaluate and .opencode/skills/ebench-evaluate in your project.
Going by SKILL.md and its folder, Ebench Evaluate needs the command-line tools its instructions call (jq and bash).
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Ebench Evaluate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Ebench Evaluate: Policy Monitor (anthropics/claude-for-legal, 9.6k stars), Cosmos Policy Evaluation (Orchestra-Research/AI-Research-SKILLs, 13k stars), Policy Monitor (anthropics/claude-for-legal, 9.6k stars) and Implementing Policy As Code With Open Policy Agent (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
InternRobotics (a GitHub organization) maintains it in InternRobotics/EBench, which has 145 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 24, 2026.
Source: InternRobotics/EBench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.