Agent skill

Convert Dataset

by AgibotTech in AgibotTech/genie_sim

Convert robot trajectory datasets between formats — currently agibot v1 → LeRobot v2.1 (parquet + HEVC/PNG-encoded MP4).

MPL-2.0Auto-check: notesData & Analytics

Install Convert Dataset

skills CLI
$ npx skills add AgibotTech/genie_sim --skill convert-dataset -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install AgibotTech/genie_sim convert-dataset --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/AgibotTech/genie_sim.git skills-src && mkdir -p .claude/skills && cp -r skills-src/source/geniesim_benchmark/skills/convert-dataset .claude/skills/convert-dataset && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
convert-dataset
GitHub stars
1.4k
Token cost
~1.5k tokens
SKILL.md length
427 words
Files
1
Skills in repo
26
Repo updated
First seen
Licence
MPL-2.0

At a glance

Convert robot trajectory datasets between formats — currently agibot v1 → LeRobot v2.1 (parquet + HEVC/PNG-encoded MP4).

  • Asks to convert agibot to lerobot
  • SKILL.md covers When to Use, Prerequisites, Workflow and Programmatic use, plus 3 more sections
  • Calls python3, apt and brew
  • Convert dataset

What it does

Convert Dataset is an agent skill from AgibotTech/genie_sim. Convert robot trajectory datasets between formats — currently agibot v1 → LeRobot v2.1 (parquet + HEVC/PNG-encoded MP4). Uses the geniesim dataset convert agibot-to-lerobot CLI verb, which wraps the geniesimbenchmark.dataset.convert.agibottolerobot Python API. Trigger: When the user asks to "convert agibot to lerobot", "convert dataset", "transcode trajectory data", "build a LeRobot dataset", "把 agibot 数据转成 lerobot", or provides an agibot episode dir / batch dir and wants the LeRobot v2.1 layout…

Its SKILL.md is about 1.5k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering DataFrames. It works with Python and FFmpeg. The repository describes itself as: Simulation Platform from AgiBot. The licence is MPL-2.0.

When your agent uses it

  • Asks to convert agibot to lerobot
  • Convert dataset
  • Transcode trajectory data
  • Build a LeRobot dataset

Example prompts

  • “convert agibot to lerobot”
  • “convert dataset”
  • “transcode trajectory data”
  • “/convert-dataset”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 6ca11c7. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python3
    • apt
    • brew
    • ffmpeg

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Convert Dataset loads about 1.5k tokens when it runs. Until then it costs about 145 tokens; SKILL.md has 427 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~145
When it runs · the whole SKILL.md, loaded when a task matches
~1.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check: notes

The automated check noted patterns worth knowing about, such as sudo or a known installer.

  • NoteRuns commands with sudoSKILL.md:55
    missing it surfaces the install hint (`sudo apt install ffmpeg` on

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from AgibotTech/genie_sim at commit 6ca11c7, republished under its MPL-2.0 licence (© AgibotTech). 427 words, ~1,539 tokens.

Download SKILL.mdSave it as .claude/skills/convert-dataset/SKILL.md (or your agent's skills folder).
name
convert-dataset
description
Convert robot trajectory datasets between formats — currently agibot v1 → LeRobot v2.1 (parquet + HEVC/PNG-encoded MP4). Uses the `geniesim dataset convert agibot-to-lerobot` CLI verb, which wraps the `geniesim_benchmark.dataset.convert.agibot_to_lerobot` Python API. Trigger: When the user asks to "convert agibot to lerobot", "convert dataset", "transcode trajectory data", "build a LeRobot dataset", "把 agibot 数据转成 lerobot", or provides an agibot episode dir / batch dir and wants the LeRobot v2.1 layout (`data/chunk-*/*.parquet` + `videos/…/*.mp4` + `meta/`).
license
MPL-2.0
metadata.author
genie-sim
metadata.version
1.0
prerequisites
geniesim_cli:fresh-machine-setup

When to Use

  • User has agibot v1 trajectory data and wants the LeRobot v2.1 layout (e.g. to feed an upstream LeRobot training pipeline, or compare against an existing LeRobot reference).
  • User provides a parent dir of multiple episode subdirs — the converter auto-detects single vs batch from layout.

Do not use for:

  • Just running a benchmark task → run-benchmark skill.
  • Probing an inference server → check-inference skill.

Prerequisites

  • geniesim_benchmark installed (tier-1 peer — comes with geniesim bootstrap).
  • ffmpeg on PATH. Used for both RGB encoding (HEVC / libx265) and depth encoding (PNG / gray16le). The converter pre-flights ffmpeg; if missing it surfaces the install hint (sudo apt install ffmpeg on Debian/Ubuntu, brew install ffmpeg on macOS).
  • h5py, numpy, pyarrow are declared deps of geniesim_benchmark; nothing to install separately.

Workflow

Single episode
bash
geniesim dataset convert agibot-to-lerobot \
  --agibot-dir ./agibot/episode_000 \
  --output-dir ./lerobot_out

--agibot-dir is treated as a single episode iff it contains aligned_joints.h5 directly. The resulting dataset has total_episodes = 1.

Batch (auto-detect)
bash
geniesim dataset convert agibot-to-lerobot \
  --agibot-dir ./agibot \
  --output-dir ./lerobot_out

When --agibot-dir does not contain aligned_joints.h5 directly, the converter scans for episode subdirectories (each must contain aligned_joints.h5). Episodes are indexed in sorted order of their directory name.

With a reference LeRobot dataset
bash
geniesim dataset convert agibot-to-lerobot \
  --agibot-dir ./agibot \
  --output-dir ./lerobot_out \
  --lerobot-ref-dir /path/to/reference/lerobot_dataset

When the agibot episode is missing the fisheye / head_back extrinsics (common — those cameras aren't on every rig), the converter pulls the missing columns from <lerobot-ref-dir>/data/chunk-000/episode_000000.parquet. Omit --lerobot-ref-dir to leave those columns empty.

Tune FPS
bash
--fps 60   # default is 30

--fps is passed to ffmpeg (-r, -framerate) and baked into the v2.1 timestamps (frame_index / fps). The meta/info.json always records fps: 30 regardless — match this if you need consistency across a collection.

Show full SKILL.md (185 more words)Show less

Programmatic use

The same conversion is callable from Python:

python
from pathlib import Path
from geniesim_benchmark.dataset.convert.agibot_to_lerobot import convert_agibot_to_lerobot

manifest = convert_agibot_to_lerobot(
    agibot_dir=Path("./agibot"),
    output_dir=Path("./lerobot_out"),
    lerobot_ref_dir=Path("./ref_lerobot"),  # optional
    fps=30.0,
)
print(manifest["total_episodes"], manifest["total_frames"])

The Python API raises RuntimeError for missing ffmpeg, missing heavy deps, or no detected episodes. The CLI wrapper catches those and prints the error to stderr with exit code 1.

Verify it worked

bash
ls -R lerobot_out/
# → data/chunk-000/episode_000000.parquet, ...
# → videos/chunk-000/{top_head,hand_left,hand_right,top_head_depth,...}/episode_*.mp4
# → meta/{info.json,tasks.jsonl,episodes.jsonl,episodes_stats.jsonl}

python3 -c "
import pyarrow.parquet as pq
t = pq.read_table('lerobot_out/data/chunk-000/episode_000000.parquet')
print(t.schema)
print('rows:', t.num_rows)
"

observation.state must be a fixed_size_list<float32, 159> and action a fixed_size_list<float32, 40> — those widths are part of the v2.1 contract and the converter writes them literally.

Troubleshooting

  • ffmpeg is not on PATH — install ffmpeg; see Prerequisites.
  • No episode directories found — --agibot-dir neither contains aligned_joints.h5 directly nor has any subdir containing one. Re-check the path; common mistake is pointing at a parent that's one level too high.
  • ERROR encoding <key>: … — ffmpeg printed something to stderr. Common causes: missing input frames (camera/<N>/<stem>.jpg glob is sparse), unsupported codec (older ffmpeg without libx265 — install ffmpeg with HEVC support, e.g. the nasm/libx265 variant), or write permission errors on --output-dir.
  • Stats look wrong — episodes_stats.jsonl reads back the parquet rows; if the parquet wasn't written the stats entry is {}. Inspect the parquet first.

Resources

© AgibotTech, MPL-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in source/geniesim_benchmark/skills/convert-dataset of AgibotTech/genie_sim.

Open the folder on GitHubat commit 6ca11c7

Compare with similar skills

Convert Dataset next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Convert Dataset compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Convert Dataset this skillAgibotTech/genie_sim1.4k—~1.5kAutomated safety check: NotesMPL-2.0
Multimodal Media Feature ExtractionTyrealQ/q-skills108—~2kAutomated safety check: NotesMIT
Chdb Datastorevemetric/vemetric3942 repos~1.4kAutomated safety check: PassApache-2.0
Polar Python SDKpolarsource/polar10k—~1.8kAutomated safety check: PassApache-2.0
CSV Data Summarizercoffeefuelbump/csv-data-summarizer-claude-skill4682 repos~1.4kAutomated safety check: PassNone
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT

Similar skills

  • Extracts pixel, video-frame, speech, music and visual-semantic features from image, video and audio files for research datasets, using local tools or the Gemini API.

    108 GitHub stars~2k tokensUpdated 14 days ago
    Data & AnalyticsAuto-check: notes
  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    394 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Polar Python SDK

    polarsource/polar

    Integrate Polar billing in server-side Python applications using the versioned Polar and PolarAsync clients.

    10k GitHub stars~1.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Retentioneering Product Analytics

    retentioneering/retentioneering-tools

    Analyze event logs, clickstreams, user paths, product funnels, retention, behavioral segments, transition graphs, step matrices, sequence patterns, and customer journeys using Retentioneering.

    920 GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed

More from AgibotTech/genie_sim

All 26 skills in this repo
  • Challenge Baseline Model

    AgibotTech/genie_sim

    Provision and launch the Simulation Challenge baseline inference model end to end: clone the inference code from a given git repo/branch, download the checkpoints from ModelScope into the repo's…

    1.4k GitHub stars~2.4k tokensUpdated 1 mo ago
    Auto-check passed
  • Challenge Download Datasets

    AgibotTech/genie_sim

    Download the Simulation Challenge LeRobot v2.1 training datasets from ModelScope using ./scripts/downloaddataset.sh.

    1.4k GitHub stars~960 tokensUpdated 1 mo ago
    Auto-check passed
  • Add Robot

    AgibotTech/genie_sim

    Bring a custom robot into the Genie Sim RT Engine — author / fix a xacro / URDF in geniesimrobotmodel, prep meshes with the offline tools (normalizeobjnames.py, diagnoseurdf.py, recomputeinertia.py…

    1.4k GitHub stars~2.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Build Workspace

    AgibotTech/genie_sim

    Build the geniesimros colcon workspace inside the Genie Sim Docker container using the geniesim ros build CLI verb.

    1.4k GitHub stars~1.2k tokensUpdated 1 mo ago
    Auto-check passed
  • Challenge Inference Protocol

    AgibotTech/genie_sim

    Reference for the Simulation Challenge inference wire protocol — the exact obs (input) and action (output) message format exchanged between the gateway/genie-sim simulator and the contestant's…

    1.4k GitHub stars~1.7k tokensUpdated 1 mo ago
    Auto-check passed
  • Challenge Login

    AgibotTech/genie_sim

    A skill your agent uses when the contestant needs to obtain or refresh their Simulation Challenge JWT (CHALLENGETOKEN), or wants to inspect the current logged-in user.

    1.4k GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed

Works with

Questions about Convert Dataset

What does Convert Dataset do?

Convert robot trajectory datasets between formats — currently agibot v1 → LeRobot v2.1 (parquet + HEVC/PNG-encoded MP4). Convert Dataset is an agent skill from AgibotTech/genie_sim.1 (parquet + HEVC/PNG-encoded MP4).

When should I use Convert Dataset?

Convert Dataset fits situations like: asks to convert agibot to lerobot; convert dataset; transcode trajectory data; build a LeRobot dataset.

How do I install Convert Dataset in Claude Code?

Run `npx skills add AgibotTech/genie_sim --skill convert-dataset -a claude-code`. Or copy the skill folder (source/geniesim_benchmark/skills/convert-dataset in AgibotTech/genie_sim) into .claude/skills/convert-dataset in your project. Claude Code loads it when a task matches its description.

How do I install Convert Dataset in Codex?

Run `npx skills add AgibotTech/genie_sim --skill convert-dataset -a codex`. Or copy the skill folder (source/geniesim_benchmark/skills/convert-dataset in AgibotTech/genie_sim) into .agents/skills/convert-dataset in your project. Codex loads it when a task matches its description.

Can I use Convert Dataset in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add AgibotTech/genie_sim --skill convert-dataset -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/convert-dataset, .gemini/skills/convert-dataset, .github/skills/convert-dataset and .opencode/skills/convert-dataset in your project.

What does Convert Dataset need to run?

Going by SKILL.md and its folder, Convert Dataset needs the command-line tools its instructions call (python3, apt, brew and ffmpeg). Our summary lists: Python 3.

Does Convert Dataset access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Convert Dataset safe to install?

Our automated static check of SKILL.md found notes only (runs commands with sudo), nothing it rates as a warning. It is not a guarantee. Review the folder before installing.

What licence does Convert Dataset use?

Convert Dataset is published under the MPL-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Convert Dataset use?

About 1.5k tokens (SKILL.md is roughly 6.2k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Convert Dataset?

Skills that share tags, products or a category with Convert Dataset: Multimodal Media Feature Extraction (TyrealQ/q-skills, 108 stars), Chdb Datastore (vemetric/vemetric, 394 stars), Polar Python SDK (polarsource/polar, 10k stars) and CSV Data Summarizer (coffeefuelbump/csv-data-summarizer-claude-skill, 468 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Convert Dataset?

AgibotTech (a GitHub organization) maintains it in AgibotTech/genie_sim, which has 1,413 GitHub stars. The repository holds 26 skills in this directory. The repository was last updated on September 7, 2026.

Source: AgibotTech/genie_sim on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.