Agent skill

Ebench Integrate Policy

by InternRobotics in InternRobotics/EBench

Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.

MITAuto-check passed

Install Ebench Integrate Policy

skills CLI
$ npx skills add InternRobotics/EBench --skill ebench-integrate-policy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install InternRobotics/EBench ebench-integrate-policy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ebench-integrate-policy .claude/skills/ebench-integrate-policy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ebench-integrate-policy
GitHub stars
145
Token cost
~1.3k tokens
SKILL.md length
576 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.

  • SKILL.md covers Establish the contract, Implement lifecycle and chunking, Connect the adapter to online… and Validate the adapter
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ebench Integrate Policy is an agent skill from InternRobotics/EBench. Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.

Its SKILL.md is about 1.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Elemental Diagnosis of Generalist Mobile Manipulation Policies. The licence is MIT.

Example prompts

  • “/ebench-integrate-policy”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 355fe56. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ebench Integrate Policy loads about 1.3k tokens when it runs. Until then it costs about 45 tokens; SKILL.md has 576 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~45
When it runs · the whole SKILL.md, loaded when a task matches
~1.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from InternRobotics/EBench at commit 355fe56, republished under its MIT licence (© InternRobotics). 576 words, ~1,285 tokens.

Download SKILL.mdSave it as .claude/skills/ebench-integrate-policy/SKILL.md (or your agent's skills folder).
name
ebench-integrate-policy
description
Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.

Integrate a policy with EBench

Resolve paths from the EBench root. Read third_party/genmanip-client/src/genmanip_client/eval_client.py and the closest adapter: baselines/X-VLA/run.py, baselines/openpi/scripts/pi_eval_client_online.py, or baselines/InternVLA-A1/inference.py. Inspect the user's policy inference API and training transforms before choosing a mapping.

Establish the contract

Observations are keyed by string worker IDs; a worker's model observation is obs[wid]["obs"]. Existing adapters use:

  • video.overlook_camera_view, video.left_camera_view, video.right_camera_view;
  • state.joints, state.gripper, state.base, and optionally state.ee_pose;
  • instruction, and reset for episode-boundary handling.

Verify image type, RGB ordering, shape, resize/padding, proprioception ordering, units, and normalization against training. Do not invent missing cameras or silently replace missing input with zeros. Keep checkpoint normalization and model transforms paired.

For the current r5a/lift2 joint-position adapters, the dispatched action orders left arm (6), left gripper (2), right arm (6), right gripper (2); base_motion has 3 components. The payload explicitly declares control_type, is_rel, and base_is_rel. Verify a different robot/control mode against its server contract rather than generalizing these dimensions.

Model outputs are not interchangeable:

  • X-VLA rescales base/gripper channels, reorders joints/grippers, and sends absolute base motion.
  • OpenPI reorders joints/grippers and differences chunk-relative base predictions into per-step deltas with base_is_rel=True.
  • InternVLA-A1 has checkpoint statistics and an action_mode setting; inspect its conversion before selecting delta or absolute behavior.

Document the chosen model-to-server channel mapping, normalization, frame/units, gripper interpretation, and relative/absolute semantics. Match the checkpoint, not whichever baseline is easiest to copy.

Implement lifecycle and chunking

Use EvalClient.reset() for initial observations and step() for execution; inspect supported single-action/chunk forms in the pinned client. Keep string worker IDs consistent. Limit deployed horizon to available predictions and reset model history, cached actions, and temporal state at each episode boundary. done indicates evaluation completion, not task success; use saved result metrics for success.

Never execute the remainder of an old chunk after an episode reset. For multiple workers, keep history/chunks independent and handle worker-specific resets. If a step times out, execution may already have occurred: do not blindly resend actions. Use bounded client recovery and fresh observations, discard stale actions, and record the interruption. Close the client in finally so recordings/results are flushed.

Show full SKILL.md (240 more words)Show less

Connect the adapter to online evaluation

Expose evaluation URL, run ID, token, and worker IDs as runtime settings rather than hardcoding them. For a new online run, follow the queue-and-launch sequence in ebench-evaluate: call gmp online submit --print_endpoint, wait for readiness, and parse its returned endpoint and task_id. No evaluation endpoint is required from the user before this submission.

Only after both values are available, construct the client using the returned assignment:

python
import os
from genmanip_client import EvalClient

client = EvalClient(
    base_url=os.environ["EVAL_URL"],  # online submit response: endpoint
    run_id=os.environ["RUN_ID"],     # online submit response: task_id
    token=os.environ["TOKEN"],
    worker_ids=["0"],
)

Connect this client to the policy's reset/inference/step loop and close it in finally. The platform base URL is used for queue submission; EvalClient.base_url is the returned evaluation endpoint. Do not create/reset workers while the task is still queued, or resubmit an online task each time the adapter reconnects. Adapter implementation alone does not require submitting a live task; use this flow when running the requested online evaluation.

Validate the adapter

Use a representative local observation or fixture to check preprocessing, finite action values, dimensions, channel order, inverse normalization, chunk-length boundaries, and reset behavior without a simulator. Include a known-value action conversion example that would expose swapped channels or incorrect delta semantics; shape-only assertions are insufficient.

Then run a small validation rollout when a server/checkpoint is available and evaluation is in scope. Inspect actual state/action traces before scaling. Deliver the adapter, concrete launch command, mapping description, and evidence distinguishing offline contract checks from live rollout validation. Keep model-specific code in its baseline/adapter directory rather than changing unrelated upstream submodules.

© InternRobotics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ebench-integrate-policy of InternRobotics/EBench.

Open the folder on GitHubat commit 355fe56

Compare with similar skills

Ebench Integrate Policy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ebench Integrate Policy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ebench Integrate Policy this skillInternRobotics/EBench145—~1.3kAutomated safety check: PassMIT
Implementing Policy As Code With Open Policy Agentmukul975/Anthropic-Cybersecurity-Skills34k—~2.6kAutomated safety check: NotesApache-2.0
Implementing Usb Device Control Policymukul975/Anthropic-Cybersecurity-Skills34k—~1.4kAutomated safety check: PassApache-2.0
Implementing GCP Organization Policy Constraintsmukul975/Anthropic-Cybersecurity-Skills34k—~1.8kAutomated safety check: PassApache-2.0
Implementing Container Network Policies With Calicomukul975/Anthropic-Cybersecurity-Skills34k—~680Automated safety check: PassApache-2.0
Implementing File Integrity Monitoring With Aidemukul975/Anthropic-Cybersecurity-Skills34k—~642Automated safety check: NotesApache-2.0

Similar skills

  • Implementing Policy As Code With Open Policy Agent

    mukul975/Anthropic-Cybersecurity-Skills

    Implements policy-as-code enforcement with Open Policy Agent (OPA) and Gatekeeper for Kubernetes and CI/CD pipelines, covering writing Rego policies, deploying OPA Gatekeeper as a Kubernetes…

    34k GitHub stars~2.6k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check: notes
  • Implementing Usb Device Control Policy

    mukul975/Anthropic-Cybersecurity-Skills

    Implements USB device control policies to restrict unauthorized removable media access on endpoints, preventing data exfiltration and malware introduction via USB devices.

    34k GitHub stars~1.4k tokensUpdated 1 mo ago
    SecurityAuto-check passed
  • Implementing GCP Organization Policy Constraints

    mukul975/Anthropic-Cybersecurity-Skills

    Implements GCP Organization Policy constraints via gcloud and Terraform, such as restricting external IPs, resource locations, default service accounts, and service account keys, plus dry-run…

    34k GitHub stars~1.8k tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Implementing Container Network Policies With Calico

    mukul975/Anthropic-Cybersecurity-Skills

    Uses Calico's own policy CRDs beyond the upstream Kubernetes API - GlobalNetworkPolicy, HostEndpoint, NetworkSet, policy tiers, and DNS-based egress rules - applied and audited with calicoctl.

    34k GitHub stars~680 tokensUpdated 1 mo ago
    DevOps & CloudAuto-check passed
  • Implementing File Integrity Monitoring With Aide

    mukul975/Anthropic-Cybersecurity-Skills

    Configures AIDE (Advanced Intrusion Detection Environment) for file integrity monitoring on Linux, covering baseline database creation, scheduled integrity checks via cron, change detection, and…

    34k GitHub stars~642 tokensUpdated 1 mo ago
    SecurityAuto-check: notes
  • Implementing Anti Ransomware Group Policy

    mukul975/Anthropic-Cybersecurity-Skills

    Configures Windows Group Policy Objects to block ransomware execution and lateral spread, covering AppLocker rules, Software Restriction Policies, Controlled Folder Access, attack surface reduction…

    34k GitHub stars~2.3k tokensUpdated 1 mo ago
    SecurityAuto-check passed

More from InternRobotics/EBench

  • Ebench Analyze

    InternRobotics/EBench

    Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.

    145 GitHub stars~1.1k tokensUpdated 14 days ago
    Auto-check passed
  • Ebench Evaluate

    InternRobotics/EBench

    Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.

    145 GitHub stars~1.9k tokensUpdated 14 days ago
    Auto-check passed
  • Ebench Setup

    InternRobotics/EBench

    Prepare or check an EBench evaluation environment for OpenPI, X-VLA, InternVLA-A1, or a custom policy.

    145 GitHub stars~1k tokensUpdated 14 days ago
    Auto-check passed
  • Ebench Debug

    InternRobotics/EBench

    Diagnose EBench evaluation failures, stalled workers, transport errors, invalid actions, and unexpectedly low scores using logs and episode artifacts.

    145 GitHub stars~862 tokensUpdated 14 days ago
    Auto-check passed

Questions about Ebench Integrate Policy

What does Ebench Integrate Policy do?

Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets. Ebench Integrate Policy is an agent skill from InternRobotics/EBench. Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.

How do I install Ebench Integrate Policy in Claude Code?

Run `npx skills add InternRobotics/EBench --skill ebench-integrate-policy -a claude-code`. Or copy the skill folder (skills/ebench-integrate-policy in InternRobotics/EBench) into .claude/skills/ebench-integrate-policy in your project. Claude Code loads it when a task matches its description.

How do I install Ebench Integrate Policy in Codex?

Run `npx skills add InternRobotics/EBench --skill ebench-integrate-policy -a codex`. Or copy the skill folder (skills/ebench-integrate-policy in InternRobotics/EBench) into .agents/skills/ebench-integrate-policy in your project. Codex loads it when a task matches its description.

Can I use Ebench Integrate Policy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add InternRobotics/EBench --skill ebench-integrate-policy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ebench-integrate-policy, .gemini/skills/ebench-integrate-policy, .github/skills/ebench-integrate-policy and .opencode/skills/ebench-integrate-policy in your project.

What does Ebench Integrate Policy need to run?

SKILL.md names no scripts, command-line tools or credentials: Ebench Integrate Policy is instructions for the agent only. Our summary lists: Python 3.

Does Ebench Integrate Policy access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ebench Integrate Policy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ebench Integrate Policy use?

Ebench Integrate Policy is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ebench Integrate Policy use?

About 1.3k tokens (SKILL.md is roughly 5.1k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ebench Integrate Policy?

Skills that share tags, products or a category with Ebench Integrate Policy: Implementing Policy As Code With Open Policy Agent (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Implementing Usb Device Control Policy (mukul975/Anthropic-Cybersecurity-Skills, 34k stars), Implementing GCP Organization Policy Constraints (mukul975/Anthropic-Cybersecurity-Skills, 34k stars) and Implementing Container Network Policies With Calico (mukul975/Anthropic-Cybersecurity-Skills, 34k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ebench Integrate Policy?

InternRobotics (a GitHub organization) maintains it in InternRobotics/EBench, which has 145 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 24, 2026.

Source: InternRobotics/EBench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.