Agent skill

Ebench Evaluate

by InternRobotics in InternRobotics/EBench

Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.

MITAuto-check passed

Install Ebench Evaluate

skills CLI
$ npx skills add InternRobotics/EBench --skill ebench-evaluate -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install InternRobotics/EBench ebench-evaluate --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ebench-evaluate .claude/skills/ebench-evaluate && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ebench-evaluate
GitHub stars
145
Token cost
~1.9k tokens
SKILL.md length
931 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.

  • Works in 4 steps: Submit and wait in the queue. gmp online… → Save the assignment. Record the returned… → Start actual model inference against the… → …
  • SKILL.md covers Define the run, Select one submission path, Online evaluation: queue,… and Launch the real policy, plus 1 more section
  • Calls jq and bash

What it does

Ebench Evaluate is an agent skill from InternRobotics/EBench. Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Elemental Diagnosis of Generalist Mobile Manipulation Policies. The licence is MIT.

Example prompts

  • “/ebench-evaluate”

Workflow steps

4 steps, taken from the first numbered list in SKILL.md.

  1. Submit and wait in the queue. gmp online submit creates the task and polls until evaluation resources are ready. Run it as a monitored…
  2. Save the assignment. Record the returned task ID and endpoint in the local run metadata, excluding credentials. Use the returned task ID…
  3. Start actual model inference against the assigned endpoint. Use the baseline commands below or the custom adapter. Pass EVAL_URL as the…
  4. Monitor until evaluation completes. Use the assigned endpoint/task ID for status and preserve the resulting logs and episode artifacts…

What it can do on your machine

Read from SKILL.md and the folder at commit 355fe56. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • jq
    • bash

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ebench Evaluate loads about 1.9k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 931 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from InternRobotics/EBench at commit 355fe56, republished under its MIT licence (© InternRobotics). 931 words, ~1,925 tokens.

Download SKILL.mdSave it as .claude/skills/ebench-evaluate/SKILL.md (or your agent's skills folder).
name
ebench-evaluate
description
Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.

Run an EBench policy evaluation

Work from the EBench root. Read the selected baseline entry point and the relevant CLI implementation under third_party/genmanip-client/src/genmanip_client/. Verify installed gmp ... --help before relying on flags from a different revision.

Define the run

Resolve model/checkpoint, track, split, server mode, GPU/worker budget, and output directory from the request. Use val_train / val_unseen for tuning. Run held-out test when requested for final evaluation; do not silently substitute a split or tune on held-out results.

Record a small manifest beside the run logs: EBench and submodule commits, local code changes, checkpoint revision/path, config and normalization source, track/split/task selection, run/task ID, worker-to-GPU mapping, action/replan horizon, start time, sanitized command, and result locations. Mark unavailable fields as unknown. Exclude tokens and signed credentials.

Select one submission path

  • Local GenManip: verify the actual server config path or benchmark alias, then use gmp submit "$CONFIG_PATH" --run_id "$RUN_ID" --host "$SERVER_HOST" --port "$SERVER_PORT". The pinned submit CLI takes host/port, not the online platform's --base_url. Submission schedules jobs; it does not load the user's policy.
  • New online task: inspect extensions/online_cli.py; gmp online submit --base_url "$PLATFORM_URL" --token "$TOKEN" --model_name "$MODEL_NAME" --benchmark_set EBench --timeout 600 --print_endpoint creates a task and waits for readiness. Choose metadata, visibility and wait budget appropriate to the user's request; 600 seconds is an example budget, not a service guarantee. Capture returned task_id and endpoint; use that task ID as the client run ID.
  • Existing online task: reuse it. Query gmp online ready --base_url "$PLATFORM_URL" --token "$TOKEN" --task_id "$RUN_ID"; use its returned evaluation endpoint. A wait timeout does not prove creation failed. Inspect existing task state before retrying creation; do not create duplicate tasks while resources are pending.

The platform URL, returned evaluation endpoint, and an OpenPI model server address serve different purposes. Do not interchange them. Preserve credentials through local environment/configuration, omit them from reports, and avoid shell tracing of authenticated commands.

Online evaluation: queue, obtain endpoint, then evaluate

For a new online evaluation, the agent should complete this sequence rather than require the user to obtain an endpoint manually. First verify the local model environment/checkpoint and obtain the platform URL and locally configured token. A new task does not need a pre-existing evaluation URL or task ID.

  1. Submit and wait in the queue. gmp online submit creates the task and polls until evaluation resources are ready. Run it as a monitored process; while pending, report that the task is waiting, not evaluating. This Bash example requires jq and uses a configurable wait budget:

    bash
    set -euo pipefail
    : "${PLATFORM_URL:?Set the online platform URL}"
    : "${TOKEN:?Set the API token locally}"
    : "${MODEL_NAME:?Set the model name}"
    READY_JSON=$(gmp online submit \
      --base_url "$PLATFORM_URL" \
      --token "$TOKEN" \
      --model_name "$MODEL_NAME" \
      --model_type VLA \
      --benchmark_set EBench \
      --timeout "${QUEUE_TIMEOUT_SECONDS:-600}" \
      --print_endpoint)
    
    # Parse only a successful ready response; never launch with empty values.
    EVAL_URL=$(printf '%s' "$READY_JSON" | jq -er '.endpoint | strings | select(length > 0)')
    RUN_ID=$(printf '%s' "$READY_JSON" | jq -er '.task_id | strings | select(length > 0)')
    export EVAL_URL RUN_ID

    --print_endpoint returns a JSON object containing both endpoint and task_id, not a plain URL. If submission fails, times out, or either field is missing, stop before launching the client. A timeout may leave a task queued: recover its ID from available logs/platform state and query gmp online ready instead of submitting again. If its ID cannot be determined, report that uncertainty rather than create a duplicate.

  2. Save the assignment. Record the returned task ID and endpoint in the local run metadata, excluding credentials. Use the returned task ID unchanged as RUN_ID; do not substitute a friendly experiment name. Keep any credential-bearing endpoint out of shared reports.

  3. Start actual model inference against the assigned endpoint. Use the baseline commands below or the custom adapter. Pass EVAL_URL as the evaluation server address and RUN_ID as the run ID. Do not run a second gmp submit against the online endpoint: the online task already schedules the evaluation. For OpenPI, ensure its separate local model server is ready before launching the eval client.

  4. Monitor until evaluation completes. Use the assigned endpoint/task ID for status and preserve the resulting logs and episode artifacts. Queue readiness only means resources are available; it does not mean the model has been evaluated.

An existing ready task starts at step 2; an existing queued task uses gmp online ready until ready within the chosen wait budget. Once both fields are valid, continue to evaluation within the user's request without asking them to copy the values back manually.

Show full SKILL.md (268 more words)Show less

Launch the real policy

gmp eval supplies fake actions. Use it only for an explicitly scoped connectivity smoke test, never as evidence of a checkpoint's performance.

For X-VLA, the existing wrapper accepts environment variables (here EVAL_URL is the evaluation endpoint):

bash
MODEL_PATH="$CHECKPOINT" BASE_URL="$EVAL_URL" RUN_ID="$RUN_ID" \
TOKEN="$TOKEN" WORKER_IDS=0 GPU_IDS=0 LOG_DIR="$RUN_LOG_DIR" \
  bash scripts/run_xvla_eval.sh

For OpenPI, use the baseline README plus the actual serve_policy.py and pi_eval_client_online.py arguments; resolve the launch template's placeholders first. The current client constructs its policy adapter for worker_ids[0], so launch one client process per worker rather than assuming one process drives all listed workers.

For InternVLA-A1, run inference.py from its baseline directory with its supported --ckpt_path, --url, --run_id, --token, and --worker_ids arguments. Check the implementation before relying on a wrapper mentioned only in documentation.

Across hosts, share the run ID and assign disjoint worker IDs. A GPU index is not a worker ID. Start with a small validation smoke run for a newly integrated policy before expanding to the requested budget; do not submit an extra held-out smoke run by default.

Monitor and finish

Use gmp status --url "$EVAL_URL" --run_id "$RUN_ID" --token "$TOKEN" with worker logs. Distinguish waiting for resources, model loading, active steps, completed episodes, and transport failures. Bound polling/recovery to the user's time budget. Do not overwrite runs, clean results, or restart unrelated workers as routine recovery.

At completion, verify server status and expected versus saved episode coverage, not just process exit code. Capture failed/missing episodes and interruptions. Locate actual client outputs (default client_results, overridable by GENMANIP_RESULT_DIR) and server outputs separately. Report completion or partial completion, run ID, manifest/log/result paths, and any blocker. A request to run evaluation does not by itself request separate leaderboard publication.

© InternRobotics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ebench-evaluate of InternRobotics/EBench.

Open the folder on GitHubat commit 355fe56

Compare with similar skills

Ebench Evaluate next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ebench Evaluate compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ebench Evaluate this skillInternRobotics/EBench145—~1.9kAutomated safety check: PassMIT
Policy Monitoranthropics/claude-for-legal9.6k3 repos~3.7kAutomated safety check: PassApache-2.0
Cosmos Policy EvaluationOrchestra-Research/AI-Research-SKILLs13k—~3.7kAutomated safety check: PassMIT
Policy Monitoranthropics/claude-for-legal9.6k2 repos~4.5kAutomated safety check: PassApache-2.0
Implementing Policy As Code With Open Policy Agentmukul975/Anthropic-Cybersecurity-Skills34k—~2.6kAutomated safety check: NotesApache-2.0
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT

Similar skills

  • Policy Monitor

    anthropics/claude-for-legal

    Official

    Keep the AI policy current with practice — weekly sweep of saved AIAs, triage results, and vendor reviews to find policy drift, or direct query for a proposed new AI practice.

    9.6k GitHub starsUsed in 3 repos~3.7k tokens
    Legal & ComplianceAuto-check passed
  • Cosmos Policy Evaluation

    Orchestra-Research/AI-Research-SKILLs

    Sets up and runs NVIDIA Cosmos Policy evaluations on the LIBERO and RoboCasa simulators, including headless GPU rendering and inference latency profiling.

    13k GitHub stars~3.7k tokensUpdated 3 mo ago
    AI & LLM EngineeringAuto-check passed
  • Policy Monitor

    anthropics/claude-for-legal

    Official

    Keep the privacy policy current with practice. An agent skill from anthropics/claude-for-legal.

    9.6k GitHub starsUsed in 2 repos~4.5k tokens
    Legal & ComplianceAuto-check passed
  • Implementing Policy As Code With Open Policy Agent

    mukul975/Anthropic-Cybersecurity-Skills

    Implements policy-as-code enforcement with Open Policy Agent (OPA) and Gatekeeper for Kubernetes and CI/CD pipelines, covering writing Rego policies, deploying OPA Gatekeeper as a Kubernetes…

    34k GitHub stars~2.6k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes
  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed

More from InternRobotics/EBench

  • Ebench Analyze

    InternRobotics/EBench

    Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.

    145 GitHub stars~1.1k tokensUpdated 13 days ago
    Auto-check passed
  • Ebench Integrate Policy

    InternRobotics/EBench

    Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.

    145 GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Ebench Setup

    InternRobotics/EBench

    Prepare or check an EBench evaluation environment for OpenPI, X-VLA, InternVLA-A1, or a custom policy.

    145 GitHub stars~1k tokensUpdated 13 days ago
    Auto-check passed
  • Ebench Debug

    InternRobotics/EBench

    Diagnose EBench evaluation failures, stalled workers, transport errors, invalid actions, and unexpectedly low scores using logs and episode artifacts.

    145 GitHub stars~862 tokensUpdated 13 days ago
    Auto-check passed

Questions about Ebench Evaluate

What does Ebench Evaluate do?

Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records. Ebench Evaluate is an agent skill from InternRobotics/EBench. Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.

How do I install Ebench Evaluate in Claude Code?

Run `npx skills add InternRobotics/EBench --skill ebench-evaluate -a claude-code`. Or copy the skill folder (skills/ebench-evaluate in InternRobotics/EBench) into .claude/skills/ebench-evaluate in your project. Claude Code loads it when a task matches its description.

How do I install Ebench Evaluate in Codex?

Run `npx skills add InternRobotics/EBench --skill ebench-evaluate -a codex`. Or copy the skill folder (skills/ebench-evaluate in InternRobotics/EBench) into .agents/skills/ebench-evaluate in your project. Codex loads it when a task matches its description.

Can I use Ebench Evaluate in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add InternRobotics/EBench --skill ebench-evaluate -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ebench-evaluate, .gemini/skills/ebench-evaluate, .github/skills/ebench-evaluate and .opencode/skills/ebench-evaluate in your project.

What does Ebench Evaluate need to run?

Going by SKILL.md and its folder, Ebench Evaluate needs the command-line tools its instructions call (jq and bash).

Does Ebench Evaluate access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ebench Evaluate safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ebench Evaluate use?

Ebench Evaluate is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ebench Evaluate use?

About 1.9k tokens (SKILL.md is roughly 7.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ebench Evaluate?

Skills that share tags, products or a category with Ebench Evaluate: Policy Monitor (anthropics/claude-for-legal, 9.6k stars), Cosmos Policy Evaluation (Orchestra-Research/AI-Research-SKILLs, 13k stars), Policy Monitor (anthropics/claude-for-legal, 9.6k stars) and Implementing Policy As Code With Open Policy Agent (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ebench Evaluate?

InternRobotics (a GitHub organization) maintains it in InternRobotics/EBench, which has 145 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 24, 2026.

Source: InternRobotics/EBench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.