Agent skill

Physicalai Train Benchmarking A Policy

by open-edge-platform in open-edge-platform/physical-ai-studio

Benchmarks a trained Physical AI Studio policy in a simulation gym and reports success metrics.

Apache-2.0Auto-check passedDocuments & Office

Install Physicalai Train Benchmarking A Policy

skills CLI
$ npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-benchmarking-a-policy -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-edge-platform/physical-ai-studio physicalai-train-benchmarking-a-policy --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-edge-platform/physical-ai-studio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/library/physicalai-train-benchmarking-a-policy .claude/skills/physicalai-train-benchmarking-a-policy && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
physicalai-train-benchmarking-a-policy
GitHub stars
133
Token cost
~1.2k tokens
SKILL.md length
373 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Benchmarks a trained Physical AI Studio policy in a simulation gym and reports success metrics.

  • Works in 5 steps: Choose API or CLI deliberately. Use the… → Confirm the policy loads from the… → Run the full benchmark with the intended… → …
  • Running physicalai benchmark
  • SKILL.md covers Python API invocation, CLI invocation, Output and Workflow, plus 3 more sections
  • Calls uv

What it does

Physicalai Train Benchmarking A Policy is an agent skill from open-edge-platform/physical-ai-studio. Benchmarks a trained Physical AI Studio policy in a simulation gym and reports success metrics. Use when running physicalai benchmark, editing configs under library/configs/benchmark, adding or changing a Benchmark class in physicalai.benchmark, tuning rollout/episode/env settings, recording rollout videos, or interpreting results.json / results.csv.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Documents & Office, covering Product metrics and CSV and tabular files. It works with Python. The repository describes itself as: Physical AI Studio is an end-to-end framework for training robots to perform tasks through imitation learning from human demonstrations. The licence is Apache-2.0.

When your agent uses it

  • Running physicalai benchmark
  • Editing configs under library/configs/benchmark
  • Changing a Benchmark class in physicalai.benchmark
  • Tuning rollout/episode/env settings

Example prompts

  • “Use the physicalai-train-benchmarking-a-policy skill to benchmark a trained Physical AI Studio policy in a simulation gym and reports success metrics”
  • “/physicalai-train-benchmarking-a-policy”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Choose API or CLI deliberately. Use the Python API for code-level tasks; use CLI for config/docs/entry-point tasks.
  2. Confirm the policy loads from the checkpoint/export before a full sweep
  3. Run the full benchmark with the intended episode/env counts.
  4. Interpret results via BenchmarkResults/TaskResult fields; compare against a baseline checkpoint on the same config.
  5. Record videos for qualitative review when a task regresses (record_mode: failures).

What it can do on your machine

Read from SKILL.md and the folder at commit e9fb4a4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Physicalai Train Benchmarking A Policy loads about 1.2k tokens when it runs. Until then it costs about 98 tokens; SKILL.md has 373 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~98
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-edge-platform/physical-ai-studio at commit e9fb4a4, republished under its Apache-2.0 licence (© open-edge-platform). 373 words, ~1,163 tokens.

Download SKILL.mdSave it as .claude/skills/physicalai-train-benchmarking-a-policy/SKILL.md (or your agent's skills folder).
name
physicalai-train-benchmarking-a-policy
description
Benchmarks a trained Physical AI Studio policy in a simulation gym and reports success metrics. Use when running physicalai benchmark, editing configs under library/configs/benchmark, adding or changing a Benchmark class in physicalai.benchmark, tuning rollout/episode/env settings, recording rollout videos, or interpreting results.json / results.csv.
license
Apache-2.0

Benchmarking a Studio Policy

Benchmarking evaluates a trained policy by rolling it out in a gym and scoring success. Benchmark classes live in library/src/physicalai/benchmark/gyms/benchmark.py (Benchmark, PushTBenchmark, LiberoBenchmark); results types in benchmark/gyms/results.py (BenchmarkResults, TaskResult); rollout logic in library/src/physicalai/eval/rollout.py (evaluate_policy). The library supports both direct Python API use and the physicalai benchmark CLI wrapper (library/src/physicalai/cli/benchmark.py).

Python API invocation

Use this path for notebooks, tests, custom scripts, or direct library integrations.

python
from physicalai.benchmark.gyms import PushTBenchmark
from physicalai.policies import ACT

policy = ACT.load_from_checkpoint("experiments/act/version_0/checkpoints/last.ckpt")
benchmark = PushTBenchmark(num_episodes=1)
results = benchmark.evaluate(policy)
print(results.summary())
results.to_json("results/benchmark/results.json")
results.to_csv("results/benchmark/results.csv")

For exported artifacts, load the Runtime-facing model first:

python
from physicalai.benchmark.gyms import PushTBenchmark
from physicalai.inference import InferenceModel

model = InferenceModel("./exports/act_policy")
results = PushTBenchmark(num_episodes=1).evaluate(model)

CLI invocation

bash
physicalai benchmark \
  --config configs/benchmark/pusht.yaml \
  --policy physicalai.policies.ACT \
  --ckpt_path experiments/act/version_0/checkpoints/last.ckpt \
  --output_dir ./results/benchmark
  • --policy — policy class path.
  • --ckpt_path — a .ckpt or an export directory.
  • --config — a benchmark config (configs/benchmark/pusht.yaml, configs/benchmark/libero.yaml) selecting the Benchmark class and its settings.
  • --output_dir — defaults to ./results/benchmark.

Override benchmark settings on the CLI, e.g. --benchmark.num_episodes 10 --benchmark.num_envs 8.

Output

  • Prints results.summary() to stdout.
  • Writes results.json and results.csv into --output_dir.
  • Optional video via config video_dir + record_mode (all | failures | successes | none).

Workflow

  1. Choose API or CLI deliberately. Use the Python API for code-level tasks; use CLI for config/docs/entry-point tasks.
    • Done when: the selected path matches the user's requested surface area.
  2. Confirm the policy loads from the checkpoint/export before a full sweep:
    bash
    physicalai benchmark --config configs/benchmark/<suite>.yaml --policy <ClassPath> --ckpt_path <path> --benchmark.num_episodes 1
    • Done when: one episode runs end-to-end and a summary prints.
  3. Run the full benchmark with the intended episode/env counts.
    • Done when: results.json and results.csv are written and the success metric is populated.
  4. Interpret results via BenchmarkResults/TaskResult fields; compare against a baseline checkpoint on the same config.
  5. Record videos for qualitative review when a task regresses (record_mode: failures).
Show full SKILL.md (133 more words)Show less

Adding or changing a Benchmark

  1. Subclass Benchmark in benchmark/gyms/ (study PushTBenchmark / LiberoBenchmark); the gym itself comes from physicalai.gyms (pusht.py, libero.py, …).
  2. Add a matching config in library/configs/benchmark/.
  3. Add tests under library/tests/unit/benchmark/.
    • Done when: uv run --no-sync pytest tests/unit/benchmark passes and a 1-episode run succeeds.

Required checks

  • The policy runs from both a .ckpt and an export dir if both are supported paths.
  • The Python API path (Benchmark(...).evaluate(...)) and CLI wrapper agree on supported inputs for user-facing benchmark changes.
  • Success/episode metrics are populated (not zero/NaN by accident) and reproducible across runs.
  • Env/episode counts match hardware; large num_envs fits memory.
  • Heavy gym deps (e.g. libero, robocasa) are gated behind their optional extras and imported lazily.
  • physicalai-train-training-a-policy — to produce the checkpoint being benchmarked.
  • physicalai-train-exporting-and-validating — when benchmarking an exported artifact for deployment parity.

© open-edge-platform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/library/physicalai-train-benchmarking-a-policy of open-edge-platform/physical-ai-studio.

Open the folder on GitHubat commit e9fb4a4

Compare with similar skills

Physicalai Train Benchmarking A Policy next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Physicalai Train Benchmarking A Policy compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Physicalai Train Benchmarking A Policy this skillopen-edge-platform/physical-ai-studio133—~1.2kAutomated safety check: PassApache-2.0
Wjs X Increasing Followerjianshuo/claude-skills131—~1.4kAutomated safety check: PassMIT
CSV Data Summarizerzrt-ai-lab/opencode-skills287—~577Automated safety check: PassNone
Excel Spreadsheet Creation and Editinganthropics/skills180k4 repos~2.1kAutomated safety check: PassProprietary
XLSXrvdbreemen/OTGW-firmware20735 repos~2.9kAutomated safety check: PassProprietary
Instrument Data To Allotropeaws-samples/amazon-bedrock-agents-healthcare-lifesciences2742 repos~2.7kAutomated safety check: PassApache-2.0

Similar skills

  • Wjs X Increasing Follower

    jianshuo/claude-skills

    A skill your agent uses when 王建硕 wants to systematically grow his X (Twitter) followers by running numbered, A/B-testable growth experiments and tracking which ones actually work.

    131 GitHub stars~1.4k tokensUpdated 1 mo ago
    Product & Project ManagementAuto-check passed
  • CSV Data Summarizer

    zrt-ai-lab/opencode-skills

    CSV数据分析技能。使用Python和pandas分析CSV文件,生成统计摘要和快速可视化图表。当用户上传或提到CSV文件、需要分析表格数据时自动使用。

    287 GitHub stars~577 tokensUpdated 2 mo ago
    Documents & OfficeAuto-check passed
  • Official

    Creates, edits and analyzes spreadsheets (.xlsx, .xlsm, .csv, .tsv) with openpyxl and pandas, writing live formulas and recalculating to confirm zero formula errors.

    180k GitHub starsUsed in 4 repos~2.1k tokens
    Documents & OfficeAuto-check passed
  • XLSX

    rvdbreemen/OTGW-firmware

    Use this skill any time a spreadsheet file is the primary input or output.

    207 GitHub starsUsed in 35 repos~2.9k tokens
    Documents & OfficeAuto-check passed
  • Instrument Data To Allotrope

    aws-samples/amazon-bedrock-agents-healthcare-lifesciences

    Official

    Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV.

    274 GitHub starsUsed in 2 repos~2.7k tokens
    Documents & OfficeAuto-check passed
  • Excel Workbook Editor

    HKUDS/DeepTutor

    Reads, creates and edits Excel workbooks with openpyxl, including formulas, styles, charts and CSV or TSV tables, with advice on formula values openpyxl cannot compute.

    41k GitHub stars~1.9k tokensUpdated 3 days ago
    Documents & OfficeAuto-check passed

More from open-edge-platform/physical-ai-studio

  • Physicalai Train Adding A Policy

    open-edge-platform/physical-ai-studio

    Adds or modifies a Physical AI Studio policy under library/src/physicalai/policies.

    133 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Physicalai Train Exporting And Validating

    open-edge-platform/physical-ai-studio

    Exports and validates Physical AI Studio policies for Runtime deployment.

    133 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Physicalai Train Training A Policy

    open-edge-platform/physical-ai-studio

    Trains, validates, tests, and runs prediction for Physical AI Studio policies via the library Lightning stack.

    133 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Physicalai Train Working With Datasets

    open-edge-platform/physical-ai-studio

    Works with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format.

    133 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Studio Adding Robot Form UI Fields

    open-edge-platform/physical-ai-studio

    Adds a new interactive robot form UI field for plugin payload schemas.

    133 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Studio Creating A Robot Plugin

    open-edge-platform/physical-ai-studio

    Creates or modifies an external Physical AI robot plugin for Studio.

    133 GitHub stars~3.1k tokensUpdated today
    Auto-check passed

Works with

Questions about Physicalai Train Benchmarking A Policy

What does Physicalai Train Benchmarking A Policy do?

Benchmarks a trained Physical AI Studio policy in a simulation gym and reports success metrics. Physicalai Train Benchmarking A Policy is an agent skill from open-edge-platform/physical-ai-studio. Benchmarks a trained Physical AI Studio policy in a simulation gym and reports success metrics.

When should I use Physicalai Train Benchmarking A Policy?

Physicalai Train Benchmarking A Policy fits situations like: running physicalai benchmark; editing configs under library/configs/benchmark; changing a Benchmark class in physicalai.benchmark; tuning rollout/episode/env settings.

How do I install Physicalai Train Benchmarking A Policy in Claude Code?

Run `npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-benchmarking-a-policy -a claude-code`. Or copy the skill folder (skills/library/physicalai-train-benchmarking-a-policy in open-edge-platform/physical-ai-studio) into .claude/skills/physicalai-train-benchmarking-a-policy in your project. Claude Code loads it when a task matches its description.

How do I install Physicalai Train Benchmarking A Policy in Codex?

Run `npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-benchmarking-a-policy -a codex`. Or copy the skill folder (skills/library/physicalai-train-benchmarking-a-policy in open-edge-platform/physical-ai-studio) into .agents/skills/physicalai-train-benchmarking-a-policy in your project. Codex loads it when a task matches its description.

Can I use Physicalai Train Benchmarking A Policy in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-benchmarking-a-policy -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/physicalai-train-benchmarking-a-policy, .gemini/skills/physicalai-train-benchmarking-a-policy, .github/skills/physicalai-train-benchmarking-a-policy and .opencode/skills/physicalai-train-benchmarking-a-policy in your project.

What does Physicalai Train Benchmarking A Policy need to run?

Going by SKILL.md and its folder, Physicalai Train Benchmarking A Policy needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Physicalai Train Benchmarking A Policy access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Physicalai Train Benchmarking A Policy safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Physicalai Train Benchmarking A Policy use?

Physicalai Train Benchmarking A Policy is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Physicalai Train Benchmarking A Policy use?

About 1.2k tokens (SKILL.md is roughly 4.7k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Physicalai Train Benchmarking A Policy?

Skills that share tags, products or a category with Physicalai Train Benchmarking A Policy: Wjs X Increasing Follower (jianshuo/claude-skills, 131 stars), CSV Data Summarizer (zrt-ai-lab/opencode-skills, 287 stars), Excel Spreadsheet Creation and Editing (anthropics/skills, 180k stars) and XLSX (rvdbreemen/OTGW-firmware, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Physicalai Train Benchmarking A Policy?

open-edge-platform (a GitHub organization) maintains it in open-edge-platform/physical-ai-studio, which has 133 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 11, 2026.

Source: open-edge-platform/physical-ai-studio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.