Agent skill

Ebench Analyze

by InternRobotics in InternRobotics/EBench

Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.

MITAuto-check passed

Install Ebench Analyze

skills CLI
$ npx skills add InternRobotics/EBench --skill ebench-analyze -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install InternRobotics/EBench ebench-analyze --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/InternRobotics/EBench.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/ebench-analyze .claude/skills/ebench-analyze && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
ebench-analyze
GitHub stars
145
Token cost
~1.1k tokens
SKILL.md length
506 words
Files
1
Skills in repo
5
Repo updated
First seen
Licence
MIT

At a glance

Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.

  • SKILL.md covers Verify inputs before rendering, Generate the report and Interpret with the right…
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Ebench Analyze is an agent skill from InternRobotics/EBench. Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.

Its SKILL.md is about 1.1k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

The repository describes itself as: Elemental Diagnosis of Generalist Mobile Manipulation Policies. The licence is MIT.

Example prompts

  • “/ebench-analyze”

What it can do on your machine

Read from SKILL.md and the folder at commit 355fe56. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are bash).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Ebench Analyze loads about 1.1k tokens when it runs. Until then it costs about 49 tokens; SKILL.md has 506 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~49
When it runs · the whole SKILL.md, loaded when a task matches
~1.1k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from InternRobotics/EBench at commit 355fe56, republished under its MIT licence (© InternRobotics). 506 words, ~1,054 tokens.

Download SKILL.mdSave it as .claude/skills/ebench-analyze/SKILL.md (or your agent's skills folder).
name
ebench-analyze
description
Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.

Analyze EBench evaluation results

Work from the EBench root. Read third_party/genmanip-client/src/genmanip_client/extensions/analyse_cli.py and analyse.py for the pinned behavior; inspect default_cluster_map.json only when interpreting taxonomy or checking task coverage.

Verify inputs before rendering

Identify each run's model/checkpoint, benchmark revision, track, split, task set, seed/episode coverage, and completion state from its manifest and result files. Keep incomplete runs labeled. Missing episodes are not automatic successes or failures; state the coverage and denominator instead of inventing outcomes.

Default discovery scans <project_root>/saved/eval_results/<benchmark>/<run_id>. Client results commonly live under client_results/<benchmark>/<run_id> instead: pass the concrete run directory explicitly. Do not point to a single seed directory or assume EBench's root contains server outputs.

The loader prefers per-episode result_info.json, then run-level result.json, then task-level episode_result.json. These formats retain different detail. Raw result_info.json can carry metric_score needed for atomic-skill aggregation; absent metrics cannot be recovered from a total success rate. Check loaded records and parse failures before drawing conclusions.

Generate the report

With actual run paths already verified:

bash
gmp analyse "$RUN_A_DIR" "$RUN_B_DIR" --no-reference -o "$REPORT_PATH"

Omit --no-reference when bundled baseline comparison is desired. The default includes bundled reference models; --reference renders only those reference data and skips local runs. Label their bundled version rather than claiming they are freshly measured or current leaderboard standings.

An HTML file is not proof that local results loaded. The pinned CLI falls back to bundled reference data when no runs/records load, even when --no-reference was supplied. Check the CLI's loaded-record messages and report payload/run IDs against the requested inputs. If empty, report missing data; never describe the fallback as the user's model performance.

Use --group 'Label=pattern' only to combine intended compatible runs. Grouping different checkpoints, splits or overlapping retries can conceal variation or double-count evidence. Explicitly identify grouping members. Choose a fresh output path so prior reports remain available.

Show full SKILL.md (218 more words)Show less

Interpret with the right denominator

  • Report SR and score separately, per split, with coverage. In the current aggregator, top-line means are over loaded records; cluster summaries first average within each task and then across tasks. Run-level aggregated input is not equivalent to raw per-episode input for weighting or uncertainty.
  • Explain capability dimensions (Scene, Atomic Skill, Horizon, Precision, Mobility) and generalization dimensions (Object, Background, Instruction, Mixed) only where task labels and metrics support them. Mark absent axes as unavailable, not zero.
  • Compare models on aligned benchmark revisions, split/task coverage, and evaluation settings. If coverage differs, present that difference and, when raw data permits, a clearly labeled matched subset; do not silently compare partial scores as a full benchmark.
  • Do not present standard deviation across task/episode records as a confidence interval across independent runs. Separate repeated-seed variation from variation between tasks.
  • Ground failure hypotheses in task-level metrics and representative episode traces. Aggregate scores alone cannot establish a camera bug, planning failure, or causal explanation. Use validation splits for suggested tuning; keep held-out results for final assessment.

Deliver a linked HTML report, input run paths/IDs, actual loaded coverage, aggregation/reference settings, key supported findings, and limitations. If only reference data or partial results exist, make that the main conclusion. Do not publish results to a leaderboard as a side effect of analysis.

© InternRobotics, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/ebench-analyze of InternRobotics/EBench.

Open the folder on GitHubat commit 355fe56

Compare with similar skills

Ebench Analyze next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Ebench Analyze compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Ebench Analyze this skillInternRobotics/EBench145—~1.1kAutomated safety check: PassMIT
Arize Evaluatorgithub/awesome-copilot40k2 repos~8.1kAutomated safety check: NotesMIT
LLM Evaluationdavila7/claude-code-templates32k13 repos~3.5kAutomated safety check: PassMIT
Agent Evaluation Reportingsickn33/agentic-awesome-skills47k1 repos~2.1kAutomated safety check: PassMIT
Agent Evaluationsickn33/agentic-awesome-skills47k1 repos~2kAutomated safety check: PassMIT
EvaluatorsArize-ai/phoenix12k—~1.7kAutomated safety check: PassCustom licence

Similar skills

  • Arize Evaluator

    github/awesome-copilot

    Official

    Handles LLM-as-judge evaluation workflows on Arize including creating/updating evaluators, running evaluations on spans or experiments, managing tasks, trigger-run operations, column mapping, and…

    40k GitHub starsUsed in 2 repos~8.1k tokens
    AI & LLM EngineeringAuto-check: notes
  • LLM Evaluation

    davila7/claude-code-templates

    Master comprehensive evaluation strategies for LLM applications, from automated metrics to human evaluation and A/B testing.

    32k GitHub starsUsed in 13 repos~3.5k tokens
    AI & LLM EngineeringAuto-check passed
  • Agent Evaluation Reporting

    sickn33/agentic-awesome-skills

    A skill your agent uses when summarizing agent evaluations where autonomous, assisted, failed, timed-out, or invalid outcomes must remain distinct and comparable.

    47k GitHub starsUsed in 1 repo~2.1k tokens
    Agent WorkflowsAuto-check passed
  • Agent Evaluation

    sickn33/agentic-awesome-skills

    Evaluate agent behavior with versioned cases and explicit verifiers.

    47k GitHub starsUsed in 1 repo~2k tokens
    Agent WorkflowsAuto-check passed
  • Evaluators

    Arize-ai/phoenix

    Author or refine a Phoenix evaluator — code or LLM-as-a-judge — that scores a run's output.

    12k GitHub stars~1.7k tokensUpdated today
    EducationAuto-check passed
  • Official

    Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified.

    40k GitHub stars~6.7k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from InternRobotics/EBench

  • Ebench Evaluate

    InternRobotics/EBench

    Run and monitor an EBench policy evaluation against a local GenManip server or the online service, including baseline launch commands, worker allocation, and reproducible run records.

    145 GitHub stars~1.9k tokensUpdated 13 days ago
    Auto-check passed
  • Ebench Integrate Policy

    InternRobotics/EBench

    Implement or review a custom VLA policy adapter for EBench EvalClient, including observation preprocessing, action semantics, chunking, and episode resets.

    145 GitHub stars~1.3k tokensUpdated 13 days ago
    Auto-check passed
  • Ebench Setup

    InternRobotics/EBench

    Prepare or check an EBench evaluation environment for OpenPI, X-VLA, InternVLA-A1, or a custom policy.

    145 GitHub stars~1k tokensUpdated 13 days ago
    Auto-check passed
  • Ebench Debug

    InternRobotics/EBench

    Diagnose EBench evaluation failures, stalled workers, transport errors, invalid actions, and unexpectedly low scores using logs and episode artifacts.

    145 GitHub stars~862 tokensUpdated 13 days ago
    Auto-check passed

Questions about Ebench Analyze

What does Ebench Analyze do?

Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics. Ebench Analyze is an agent skill from InternRobotics/EBench. Generate and interpret EBench evaluation reports, compare runs and baselines, and diagnose capability or generalization gaps with explicit data coverage and aggregation semantics.

How do I install Ebench Analyze in Claude Code?

Run `npx skills add InternRobotics/EBench --skill ebench-analyze -a claude-code`. Or copy the skill folder (skills/ebench-analyze in InternRobotics/EBench) into .claude/skills/ebench-analyze in your project. Claude Code loads it when a task matches its description.

How do I install Ebench Analyze in Codex?

Run `npx skills add InternRobotics/EBench --skill ebench-analyze -a codex`. Or copy the skill folder (skills/ebench-analyze in InternRobotics/EBench) into .agents/skills/ebench-analyze in your project. Codex loads it when a task matches its description.

Can I use Ebench Analyze in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add InternRobotics/EBench --skill ebench-analyze -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/ebench-analyze, .gemini/skills/ebench-analyze, .github/skills/ebench-analyze and .opencode/skills/ebench-analyze in your project.

What does Ebench Analyze need to run?

SKILL.md names no scripts, command-line tools or credentials: Ebench Analyze is instructions for the agent only.

Does Ebench Analyze access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Ebench Analyze safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Ebench Analyze use?

Ebench Analyze is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Ebench Analyze use?

About 1.1k tokens (SKILL.md is roughly 4.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Ebench Analyze?

Skills that share tags, products or a category with Ebench Analyze: Arize Evaluator (github/awesome-copilot, 40k stars), LLM Evaluation (davila7/claude-code-templates, 32k stars), Agent Evaluation Reporting (sickn33/agentic-awesome-skills, 47k stars) and Agent Evaluation (sickn33/agentic-awesome-skills, 47k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Ebench Analyze?

InternRobotics (a GitHub organization) maintains it in InternRobotics/EBench, which has 145 GitHub stars. The repository holds 5 skills in this directory. The repository was last updated on September 24, 2026.

Source: InternRobotics/EBench on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.