Agent skill

Physicalai Train Working With Datasets

by open-edge-platform in open-edge-platform/physical-ai-studio

Works with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format.

Apache-2.0Auto-check passedDevelopment

Install Physicalai Train Working With Datasets

skills CLI
$ npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-working-with-datasets -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-edge-platform/physical-ai-studio physicalai-train-working-with-datasets --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-edge-platform/physical-ai-studio.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/library/physicalai-train-working-with-datasets .claude/skills/physicalai-train-working-with-datasets && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
physicalai-train-working-with-datasets
GitHub stars
131
Token cost
~1.2k tokens
SKILL.md length
430 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Works with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format.

  • Works in 5 steps: Pick the dataset by repo_id and confirm… → Verify a batch through the Python API… → Verify CLI parity when the dataset is… → …
  • Wiring physicalai.data.lerobot.LeRobotDataModule into a training config
  • SKILL.md covers Python API usage, Wiring data into a training…, Workflow and Debugging dataloading, plus 3 more sections
  • Calls uv

What it does

Physicalai Train Working With Datasets is an agent skill from open-edge-platform/physical-ai-studio. Works with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format. Use when wiring physicalai.data.lerobot.LeRobotDataModule into a training config, choosing a repoid, converting between the physicalai and lerobot data layouts, defining observation Features/FeatureType, setting normalization, or debugging batch shapes and dataloading.

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Development, covering Database schema design and Debugging. It works with Python. The repository describes itself as: Physical AI Studio is an end-to-end framework for training robots to perform tasks through imitation learning from human demonstrations. The licence is Apache-2.0.

When your agent uses it

  • Wiring physicalai.data.lerobot.LeRobotDataModule into a training config
  • Choosing a repoid
  • Converting between the physicalai and lerobot data layouts
  • Defining observation Features/FeatureType

Example prompts

  • “Use the physicalai-train-working-with-datasets skill to work with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format”
  • “/physicalai-train-working-with-datasets”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the first numbered list in SKILL.md.

  1. Pick the dataset by repo_id and confirm its features (image keys, state dim, action dim) match the target policy's Config.
  2. Verify a batch through the Python API before training
  3. Verify CLI parity when the dataset is configured through YAML
  4. Convert layouts only when needed via converters.py (DataFormat.physicalai ↔ DataFormat.lerobot); keep field names stable, since they…
  5. Set normalization through NormalizationParameters/Feature consistently with what the policy expects at inference.

What it can do on your machine

Read from SKILL.md and the folder at commit 429ffd4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Physicalai Train Working With Datasets loads about 1.2k tokens when it runs. Until then it costs about 101 tokens; SKILL.md has 430 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~101
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-edge-platform/physical-ai-studio at commit 429ffd4, republished under its Apache-2.0 licence (© open-edge-platform). 430 words, ~1,201 tokens.

Download SKILL.mdSave it as .claude/skills/physicalai-train-working-with-datasets/SKILL.md (or your agent's skills folder).
name
physicalai-train-working-with-datasets
description
Works with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format. Use when wiring physicalai.data.lerobot.LeRobotDataModule into a training config, choosing a repo_id, converting between the physicalai and lerobot data layouts, defining observation Features/FeatureType, setting normalization, or debugging batch shapes and dataloading.
license
Apache-2.0

Working with Studio Datasets

Studio data lives in library/src/physicalai/data/. Datasets use the LeRobot format and are consumed through Lightning datamodules. The datamodules are first-class Python API objects; YAML/CLI configs are a serialization of the same construction path.

Key modules:

  • data/lerobot/datamodule.py — LeRobotDataModule (the class configs reference as physicalai.data.lerobot.LeRobotDataModule).
  • data/lerobot/dataset.py — LeRobot dataset wrapper.
  • data/lerobot/converters.py — DataFormat (StrEnum: physicalai, lerobot) and bidirectional field mapping between the two layouts.
  • data/observation.py — Observation, Feature, FeatureType, NormalizationParameters.
  • data/datamodules.py — base DataModule (Lightning LightningDataModule, auto num-workers heuristic).
  • data/dataset.py — base Dataset; data/gym.py — GymDataset for gym-generated data.

Python API usage

Use this path for notebooks, tests, direct batch inspection, or debugging dataloading without involving the training CLI.

python
from physicalai.data import LeRobotDataModule

datamodule = LeRobotDataModule(repo_id="lerobot/pusht", train_batch_size=2)
datamodule.prepare_data()
datamodule.setup("fit")
batch = next(iter(datamodule.train_dataloader()))

Done when: the batch contains the observation/action fields the policy expects, with the expected batch/action dimensions.

Wiring data into a training config

In a physicalai fit config, the data block selects the datamodule and its repo_id:

yaml
data:
  class_path: physicalai.data.lerobot.LeRobotDataModule
  init_args:
    repo_id: lerobot/pusht
    train_batch_size: 64

repo_id points at a LeRobot/HuggingFace dataset; the datamodule pulls it on first use. See the physicalai-train-training-a-policy skill for the full config.

Workflow

  1. Pick the dataset by repo_id and confirm its features (image keys, state dim, action dim) match the target policy's Config.
    • Done when: the policy's expected Feature names and action dimension line up with the dataset.
  2. Verify a batch through the Python API before training:
    python
    datamodule.prepare_data()
    datamodule.setup("fit")
    batch = next(iter(datamodule.train_dataloader()))
    • Done when: the batch has correct keys and shapes without invoking the CLI.
  3. Verify CLI parity when the dataset is configured through YAML:
    bash
    physicalai fit --config <config.yaml> --trainer.fast_dev_run=true
    • Done when: one batch flows through with correct shapes and no missing-feature errors.
  4. Convert layouts only when needed via converters.py (DataFormat.physicalai ↔ DataFormat.lerobot); keep field names stable, since they propagate to training and export.
  5. Set normalization through NormalizationParameters/Feature consistently with what the policy expects at inference.
Show full SKILL.md (150 more words)Show less

Debugging dataloading

  • Missing/renamed feature → the config's dataset features disagree with the policy; align Feature names in data/observation.py conventions.
  • Slow/stalled first batch → the LeRobot repo_id is downloading; expected on first run (see the requires_download test marker for tests that need this).
  • Wrong batch dimensions → check train_batch_size and the datamodule's collate/observation handling before changing the policy.
  • OOM or heavy swapping during training (common on smaller policies like ACT/SmolVLA on low-RAM machines) → try pin_memory=False and/or persistent_workers=False on the DataModule; see library/docs/explanation/data/datamodules.md.

Required checks

  • Feature names, FeatureType, action dim, and normalization match between dataset, Config, and any export metadata.
  • Conversions round-trip without dropping or renaming fields.
  • Direct datamodule API construction and YAML config construction produce compatible batches.
  • Tests that require downloads are marked requires_download; keep default uv run --no-sync pytest runnable offline.

Verify

bash
# from library/
uv run --no-sync pytest tests/unit/data tests/unit/datamodules
  • physicalai-train-training-a-policy — the data block is one half of a training config.
  • physicalai-train-adding-a-policy — align observation features with the policy Config.

© open-edge-platform, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/library/physicalai-train-working-with-datasets of open-edge-platform/physical-ai-studio.

Open the folder on GitHubat commit 429ffd4

Compare with similar skills

Physicalai Train Working With Datasets next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Physicalai Train Working With Datasets compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Physicalai Train Working With Datasets this skillopen-edge-platform/physical-ai-studio131—~1.2kAutomated safety check: PassApache-2.0
LangBot Plugin Developmentlangbot-app/LangBot18k—~3.9kAutomated safety check: PassApache-2.0
Python Performance Optimizationwshobson/agents40k13 repos~814Automated safety check: PassMIT
Git History Bug Auditben-manes/caffeine18k—~3.3kAutomated safety check: PassApache-2.0
Keybase RPC Log Analysiskeybase/client9.3k—~3kAutomated safety check: PassBSD-3-Clause
The Art of Debuggingstas00/the-art-of-debugging1.7k—~6.1kAutomated safety check: NotesCC-BY-SA-4.0

Similar skills

  • LangBot Plugin Development

    langbot-app/LangBot

    Guides building, debugging and testing LangBot plugins: components, SDK calls, README and locale rules, SDK pitfalls and WebSocket-based testing.

    18k GitHub stars~3.9k tokensUpdated today
    DevelopmentAuto-check passed
  • Profiles slow Python code with cProfile and memory profilers, then applies targeted fixes for CPU, memory, I/O and query bottlenecks.

    40k GitHub starsUsed in 13 repos~814 tokens
    DevelopmentAuto-check passed
  • Git History Bug Audit

    ben-manes/caffeine

    Audits a module by walking its git history commit by commit, tracking unresolved issues forward, and reporting the ones that survive to HEAD as findings.

    18k GitHub stars~3.3k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Captures a clean Keybase service log and analyzes it for redundant, duplicated or looping RPCs, then checks whether a caching fix reduced the calls.

    9.3k GitHub stars~3k tokensUpdated today
    DevelopmentAuto-check passed
  • The Art of Debugging

    stas00/the-art-of-debugging

    Condensed debugging method and tool recipes for Unix, Python and PyTorch programs: crashes, hangs, segfaults, wrong output, CUDA OOM, NaN values and slowness.

    1.7k GitHub stars~6.1k tokensUpdated 3 days ago
    DevelopmentAuto-check: notes
  • Adk Setup

    google/adk-python

    Official

    Sets up a local ADK Python development environment in a git clone of the open-source adk-python repository: a uv virtual environment, all dependency extras, pre-commit hooks, and a first unit-test…

    22k GitHub stars~993 tokensUpdated today
    DevelopmentAuto-check: notes

More from open-edge-platform/physical-ai-studio

  • Physicalai Train Adding A Policy

    open-edge-platform/physical-ai-studio

    Adds or modifies a Physical AI Studio policy under library/src/physicalai/policies.

    131 GitHub stars~1.4k tokensUpdated today
    Auto-check passed
  • Physicalai Train Exporting And Validating

    open-edge-platform/physical-ai-studio

    Exports and validates Physical AI Studio policies for Runtime deployment.

    131 GitHub stars~1.1k tokensUpdated today
    Auto-check passed
  • Physicalai Train Benchmarking A Policy

    open-edge-platform/physical-ai-studio

    Benchmarks a trained Physical AI Studio policy in a simulation gym and reports success metrics.

    131 GitHub stars~1.2k tokensUpdated today
    Auto-check passed
  • Physicalai Train Training A Policy

    open-edge-platform/physical-ai-studio

    Trains, validates, tests, and runs prediction for Physical AI Studio policies via the library Lightning stack.

    131 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Studio Adding Robot Form UI Fields

    open-edge-platform/physical-ai-studio

    Adds a new interactive robot form UI field for plugin payload schemas.

    131 GitHub stars~1.6k tokensUpdated today
    Auto-check passed
  • Studio Creating A Robot Plugin

    open-edge-platform/physical-ai-studio

    Creates or modifies an external Physical AI robot plugin for Studio.

    131 GitHub stars~3.1k tokensUpdated today
    Auto-check passed

Works with

Questions about Physicalai Train Working With Datasets

What does Physicalai Train Working With Datasets do?

Works with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format. Physicalai Train Working With Datasets is an agent skill from open-edge-platform/physical-ai-studio. Works with Physical AI Studio datasets and Lightning datamodules built on the LeRobot format.

When should I use Physicalai Train Working With Datasets?

Physicalai Train Working With Datasets fits situations like: wiring physicalai.data.lerobot.LeRobotDataModule into a training config; choosing a repoid; converting between the physicalai and lerobot data layouts; defining observation Features/FeatureType.

How do I install Physicalai Train Working With Datasets in Claude Code?

Run `npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-working-with-datasets -a claude-code`. Or copy the skill folder (skills/library/physicalai-train-working-with-datasets in open-edge-platform/physical-ai-studio) into .claude/skills/physicalai-train-working-with-datasets in your project. Claude Code loads it when a task matches its description.

How do I install Physicalai Train Working With Datasets in Codex?

Run `npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-working-with-datasets -a codex`. Or copy the skill folder (skills/library/physicalai-train-working-with-datasets in open-edge-platform/physical-ai-studio) into .agents/skills/physicalai-train-working-with-datasets in your project. Codex loads it when a task matches its description.

Can I use Physicalai Train Working With Datasets in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-edge-platform/physical-ai-studio --skill physicalai-train-working-with-datasets -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/physicalai-train-working-with-datasets, .gemini/skills/physicalai-train-working-with-datasets, .github/skills/physicalai-train-working-with-datasets and .opencode/skills/physicalai-train-working-with-datasets in your project.

What does Physicalai Train Working With Datasets need to run?

Going by SKILL.md and its folder, Physicalai Train Working With Datasets needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Physicalai Train Working With Datasets access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Physicalai Train Working With Datasets safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Physicalai Train Working With Datasets use?

Physicalai Train Working With Datasets is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Physicalai Train Working With Datasets use?

About 1.2k tokens (SKILL.md is roughly 4.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Physicalai Train Working With Datasets?

Skills that share tags, products or a category with Physicalai Train Working With Datasets: LangBot Plugin Development (langbot-app/LangBot, 18k stars), Python Performance Optimization (wshobson/agents, 40k stars), Git History Bug Audit (ben-manes/caffeine, 18k stars) and Keybase RPC Log Analysis (keybase/client, 9.3k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Physicalai Train Working With Datasets?

open-edge-platform (a GitHub organization) maintains it in open-edge-platform/physical-ai-studio, which has 131 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 9, 2026.

Source: open-edge-platform/physical-ai-studio on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.