Agent skill

Datagen Reduce Dataset Snapshots

by open-thoughts in open-thoughts/OpenThoughts-Agent

Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality.

Apache-2.0Auto-check passedDevOps & Cloud

Install Datagen Reduce Dataset Snapshots

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-reduce-dataset-snapshots -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-reduce-dataset-snapshots --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/datagen-reduce-dataset-snapshots .claude/skills/datagen-reduce-dataset-snapshots && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datagen-reduce-dataset-snapshots
GitHub stars
301
Token cost
~2.7k tokens
SKILL.md length
1,242 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality.

  • Works in 6 steps: Count the current artifact. Extract the… → Diagnose the env-hash driver. Find the… → Rewrite the env logic to group (only if… → …
  • A dataset is flagged SnapshotCapExceeded / N unique environments with N over the threshold (target < 10)
  • SKILL.md covers Thresholds — set BEFORE…, Authoritative count tool, The cycle and Decision at each round (don't…, plus 2 more sections
  • Calls git, python and pip

What it does

Datagen Reduce Dataset Snapshots is an agent skill from open-thoughts/OpenThoughts-Agent. Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality. Use when a dataset is flagged "SnapshotCapExceeded" / "N unique environments" with N over the threshold (target < 10), e.g. swegym at 906. The loop: set snapshot+oracle thresholds → count → diagnose the env-hash driver → group/unionize Dockerfiles in the patcher → regenerate + upload → re-count → TWO-TIER quality gate (harbor infra smoke +…

Its SKILL.md is about 2.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in DevOps & Cloud, covering Containers and Quality gates. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • A dataset is flagged SnapshotCapExceeded / N unique environments with N over the threshold (target < 10)
  • Tasks that involve Containers
  • Tasks that involve Quality gates

Example prompts

  • “SnapshotCapExceeded”
  • “N unique environments”
  • “/datagen-reduce-dataset-snapshots”

Requirements

  • Python 3

Workflow steps

6 steps, taken from the first numbered list in SKILL.md.

  1. Count the current artifact. Extract the flagged HF dataset → count. Confirm
  2. Diagnose the env-hash driver. Find the patcher. **Most live in the shared
  3. Rewrite the env logic to group (only if step 2's fresh sample is still high).
  4. Regenerate the full dataset + upload to a NEW repo. Never overwrite the
  5. Quality gate — TWO-TIER (DO NOT SKIP). Snapshot reduction is valid only if the
  6. Record in the tracker. Add the new repo to

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • git
    • python
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git and pip, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Datagen Reduce Dataset Snapshots loads about 2.7k tokens when it runs. Until then it costs about 179 tokens; SKILL.md has 1,242 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~179
When it runs · the whole SKILL.md, loaded when a task matches
~2.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,242 words, ~2,743 tokens.

Download SKILL.mdSave it as .claude/skills/datagen-reduce-dataset-snapshots/SKILL.md (or your agent's skills folder).
name
datagen-reduce-dataset-snapshots
description
Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality. Use when a dataset is flagged "SnapshotCapExceeded" / "N unique environments" with N over the threshold (target < 10), e.g. swegym at 906. The loop: set snapshot+oracle thresholds → count → diagnose the env-hash driver → group/unionize Dockerfiles in the patcher → regenerate + upload → re-count → TWO-TIER quality gate (harbor infra smoke + `--stages oracle` gold-patch yield) → navigate the snapshot↔fidelity tradeoff within a bounded iteration budget → record. Runs LOCALLY on the Mac + Daytona (no GPU).

datagen-reduce-dataset-snapshots

Harbor's Daytona backend builds one container snapshot per unique environment directory, keyed by a content hash of environment/ — which for our patchers is just environment/Dockerfile (solution/tests/metadata live in sibling dirs and don't affect the hash). Daytona enforces a HARD org cap (40) and a per-launch max_new_snapshots (10). A dataset whose tasks each render a distinct Dockerfile explodes to ~1 snapshot/task and is unlaunchable. Fix: make the patcher render a small shared set of Dockerfiles (grouped by a coarse key like Python version); clone the repo + run repo-specific install at agent/verifier runtime instead of image-build time, so thousands of tasks collapse onto a handful of environments.

Thresholds — set BEFORE regenerating (step 0, write in the log)

Snapshot reduction is lossy: fewer envs → less each task's env is tailored → some repos' installs stop reproducing the gold patch → oracle yield drops. You are choosing an operating point on the snapshot↔fidelity curve, not "fixing a bug" — decide these up front to avoid an unbounded chase:

  • Snapshot ceiling: hard < 10 (ideally ≤ 8). Non-negotiable — it's the cap. (20 is the "skip the dataset" line; this skill pulls a dataset back under it.)
  • Oracle-yield floor: a number set up front (e.g. ≥ 80% for a dataset destined for RL/datagen verification; lower only with explicit reason). Without a pre-set floor every result looks "one more round will help" and you over-fit get_specs to the 40-task sample.
  • Iteration budget: max regenerate→oracle rounds (e.g. 2–3). Each round is a full regenerate + upload + oracle sample (slow + Daytona builds); diminishing returns set in fast once the easy repo-family fixes are in.
  • Sample size for the yield estimate: 40 is a usable read; don't re-sample endlessly chasing a ±few-point wobble — that is not signal.

Authoritative count tool

scripts/harbor/count_snapshots_from_tasks.py computes the exact content-hash dedup count Daytona's auto_snapshot path uses (get_task_environment_hash / analyze_task_dockerfiles). Run it on a local task dir (post-extraction or post-generation), not a live HF id, to skip the registry round-trip:

bash
PY=/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python
$PY -m scripts.harbor.count_snapshots_from_tasks --local-dataset <tasks_dir>
# read the "UNIQUE ENVIRONMENTS (SNAPSHOTS): N" line

For an uploaded HF dataset, extract first:

bash
$PY -m scripts.datagen.extract_tasks_from_parquet \
  --parquet <hf-id> --output_dir /tmp/snapcount-<slug> --on_exist overwrite
$PY -m scripts.harbor.count_snapshots_from_tasks --local-dataset /tmp/snapcount-<slug>

The cycle

  1. Count the current artifact. Extract the flagged HF dataset → count. Confirm it's genuinely over threshold — use the tool above, not row counts.

  2. Diagnose the env-hash driver. Find the patcher. Most live in the shared data/patchers/ dir (patch_<name>_tasks.py, patch_exp_rpt_*_tasks.py, patch_mix_h*_tasks.py, patch_code_contests_tasks.py, …); only a few datasets keep a per-source-subdir patcher (data/<name>/generate_patched.py, e.g. swegym, swesmith). ls data/patchers/ | grep -i <name>; ls -d data/<name> 2>/dev/null. The unique-env count == number of distinct rendered Dockerfile strings. The explosion almost always comes from per-task interpolation into the Dockerfile: repo@commit in a build-time git clone/RUN, a per-instance base image, or per-task apt pins. Generate a small sample with the current patcher and count it — if the uploaded artifact is high but a fresh sample is low, the artifact is just stale (made by an older patcher) and step 4 is a pure regenerate+reupload.

  3. Rewrite the env logic to group (only if step 2's fresh sample is still high). Make the Dockerfile depend on a coarse grouping key (e.g. Python version), not the task. The swegym pattern:

    • Base image + a union of system apt deps per group; nothing task-specific in the Dockerfile.
    • Move git clone <repo>@<commit> and repo-specific pip install/make into instruction.md (agent setup), solution/solve.sh, tests/test.sh — they run at trial time, against the shared image.
    • Keep a get_specs(repo, version) map so each repo still gets its correct Python
      • install command; only the Dockerfile-affecting part is coarsened. Re-generate a sample → count → iterate the grouping until < 10.
  4. Regenerate the full dataset + upload to a NEW repo. Never overwrite the validated artifact; bump the version suffix (...-validated-v2 → ...-v3, or ...-snap-reduced). Run the patcher with --limit <=0> (no limit), --target-repo laion/<new-name> (public per feedback_hf_public_default). Then re-extract + re-count the uploaded repo to confirm < 10 end-to-end.

  5. Quality gate — TWO-TIER (DO NOT SKIP). Snapshot reduction is valid only if the tasks still build AND stay verifiable. Two distinct signals; conflating them is the classic mistake:

    • Tier 1 — infra (harbor smoke): does the env build + the agent run without crashing? Run:
      bash
      echo "laion/<new-name>" > /tmp/snap_check.md
      FORCE_COLOR=1 SAMPLE_SIZE=200 ./scripts/daytona/batch_validate_from_md.sh /tmp/snap_check.md
      # summary TSV: /Users/benjaminfeuer/Documents/agent-traces-analysis/summary.tsv
      Read infra_rate. Ignore this script's solve_rate — it's the agent's task-solve rate, low by design, NOT a measure of task well-formedness.
    • Tier 2 — oracle correctness (THE real quality gate): does the gold patch still make tests pass? batch_validate does NOT run this; run it explicitly:
      bash
      $PY scripts/daytona/validate_and_upload_from_hf.py \
        --repo_id laion/<new-name> --extract_dir <cache> \
        --stages oracle --sample_size 40 --sample_seed 42 --skip_upload \
        --keep_failed_dir <dir>/oracle_failures
      # prints "Success: S  Fail: F  Missing: M" → oracle pass = S/(S+F)
      A low oracle rate means the env/install/test harness no longer reproduces the conditions the patch needs — the env-collapse broke repo-specific installs.
    • CAP-SAFETY: do NOT re-validate the OLD high-env artifact. Sampling 200 (or
      1. tasks from a 906-env dataset tries to build that many snapshots and blasts the org cap. The new low-env artifact is cheap to validate; for a baseline use the old artifact's recorded validation number (tracker / a prior summary.tsv), or judge against the floor set in step 0 above.
    • If oracle pass < threshold, the grouping broke some repos' installs — inspect oracle_failures/ + traces/, fix the per-repo install in get_specs, regenerate, re-test (within the budget).
  6. Record in the tracker. Add the new repo to notes/RL/a3/a3_rl_tracker.md (and, for datagen rows, the MiniMax tracker experiments/active/datagen/minimax-m2.7-tt/tracker.md): the new HF id, before→after snapshot count, and smoke-test infra/solve rates. Write a dated log to /Users/benjaminfeuer/Documents/agent_logs/.

Show full SKILL.md (365 more words)Show less

Decision at each round (don't iterate reflexively)

  1. Oracle ≥ floor at < 10 snapshots → SHIP (record both numbers; done).
  2. Oracle < floor, failures cluster on a few repo-families with identifiable missing install steps → one targeted get_specs round (add the apt pkg / pip constraint / install command for those families), regenerate, re-oracle. Spend a budget slot.
  3. Oracle < floor, failures spread evenly (systemic — the shared env can't satisfy the repo diversity) → the inherent tradeoff, not a bug. Do NOT keep tweaking get_specs. Pick a coarser-but-larger grouping that still fits the cap (e.g. group by py-version × repo-family → maybe 8–15 envs instead of 5) and re-measure the (snapshots, oracle) point. Present the tradeoff curve.
  4. Budget exhausted and still < floor at any ≤-cap grouping → STOP and surface to the user with the curve (e.g. "5 envs → 48% oracle; 12 → 74%; 30 → 91% but busts the 10-cap"). Shipping a lower-fidelity dataset, raising the floor, or shelving is the user's call, not an infinite loop's.

Worked example — swegym (#31, 906 → 5 snapshots, a FAILED reduction)

laion/swegym-tasks-patched-validated-v2: 989 tasks → 906 unique envs (≈1:1). The patcher data/swegym/generate_patched.py already groups Dockerfiles by Python version (get_specs() interpolates only {python_version} + {extra_packages} from a fixed apt_map, clones the repo at runtime), so v2 was a stale artifact from an older per-task patcher. Regenerating with the current patcher → laion/swegym-tasks-patched-validated-v3 → 906 → 5 snapshots (no patcher code change). But the quality gate exposed the tradeoff: Tier-1 harbor smoke = 100% infra (looks great, would be mistaken for success); Tier-2 oracle = 19/40 = 47.5% — 5 shared py-version envs can't satisfy every repo's install. v3 is snapshots-green but oracle-red = a FAILED reduction, not shippable.

Guardrails

  • Never ship below the oracle floor just because snapshots are green. A tiny-snapshot dataset whose gold patches don't verify is worse than useless for RL/datagen (the reward signal is broken). Snapshots-green + oracle-red is a failed reduction, recorded as such.
  • Never raise/bypass the Daytona cap (max_new_snapshots, max_org_snapshots) or convert SnapshotCapExceeded to a warning — reduce the real count (feedback_daytona_snapshot_caps_hard_limit).
  • Never overwrite the existing validated artifact — always a new versioned repo.
  • Uploads are PUBLIC by default to laion/; enable_db_registration stays off (these are task datasets, not models).
  • On the Mac, run everything with the otagent python (/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python); source "${DC_AGENT_SECRET_ENV:?set DC_AGENT_SECRET_ENV first}" first (.agents/secret.md).

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/datagen-reduce-dataset-snapshots of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Datagen Reduce Dataset Snapshots next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Datagen Reduce Dataset Snapshots compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Datagen Reduce Dataset Snapshots this skillopen-thoughts/OpenThoughts-Agent301—~2.7kAutomated safety check: PassApache-2.0
Hoi Object Reconstruction Doctornvidia-isaac/video_to_data861—~1kAutomated safety check: PassCustom licence
CI CD And Automationdzhalaevd/Donatello1357 repos~2.7kAutomated safety check: NotesApache-2.0
Crabbox Quickstartopenclaw/crabbox1.5k—~1.6kAutomated safety check: NotesMIT
Ego Reconstruction Setupnvidia-isaac/video_to_data861—~826Automated safety check: PassCustom licence
Build Imageskubernetes-sigs/cloud-provider-azure294—~1.7kAutomated safety check: PassApache-2.0

Similar skills

  • Hoi Object Reconstruction Doctor

    nvidia-isaac/video_to_data

    Diagnose and repair failures in this repository's BundleSDF or SAM3D HOI object reconstruction workflow.

    861 GitHub stars~1k tokensUpdated yesterday
    Testing & QAAuto-check passed
  • CI CD And Automation

    dzhalaevd/Donatello

    Automates CI/CD pipeline setup. An agent skill from dzhalaevd/Donatello.

    135 GitHub starsUsed in 7 repos~2.7k tokens
    DevOps & CloudAuto-check: notes
  • Crabbox Quickstart

    openclaw/crabbox

    Gets you running your repository's tests in a disposable Docker or Podman container on your own machine with Crabbox, with no account and no cloud spend.

    1.5k GitHub stars~1.6k tokensUpdated today
    DevOps & CloudAuto-check: notes
  • Ego Reconstruction Setup

    nvidia-isaac/video_to_data

    Prepare the repository-local egocentric reconstruction pipeline for Codex-driven work.

    861 GitHub stars~826 tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Build Images

    kubernetes-sigs/cloud-provider-azure

    Official

    Build cloud-provider-azure container images through the repo Makefile with explicit IMAGETAG and IMAGEREGISTRY inputs, optional make flag overrides, and opt-in bounded Docker or Podman retries.

    294 GitHub stars~1.7k tokensUpdated yesterday
    DevOps & CloudAuto-check passed
  • Cb Build Test

    BlkLeg/CircuitBreaker

    How Circuit Breaker is built, tested, packaged, and kept secret-safe — the make dev/verify/test targets, the PostgreSQL integration test database and its fixtures, the mono Docker image and native…

    201 GitHub stars~1.9k tokensUpdated 4 days ago
    DevOps & CloudAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 11 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed

Questions about Datagen Reduce Dataset Snapshots

What does Datagen Reduce Dataset Snapshots do?

Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality. Datagen Reduce Dataset Snapshots is an agent skill from open-thoughts/OpenThoughts-Agent. Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality.

When should I use Datagen Reduce Dataset Snapshots?

Datagen Reduce Dataset Snapshots fits situations like: A dataset is flagged SnapshotCapExceeded / N unique environments with N over the threshold (target < 10); tasks that involve Containers; tasks that involve Quality gates.

How do I install Datagen Reduce Dataset Snapshots in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-reduce-dataset-snapshots -a claude-code`. Or copy the skill folder (.agents/skills/datagen-reduce-dataset-snapshots in open-thoughts/OpenThoughts-Agent) into .claude/skills/datagen-reduce-dataset-snapshots in your project. Claude Code loads it when a task matches its description.

How do I install Datagen Reduce Dataset Snapshots in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-reduce-dataset-snapshots -a codex`. Or copy the skill folder (.agents/skills/datagen-reduce-dataset-snapshots in open-thoughts/OpenThoughts-Agent) into .agents/skills/datagen-reduce-dataset-snapshots in your project. Codex loads it when a task matches its description.

Can I use Datagen Reduce Dataset Snapshots in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-reduce-dataset-snapshots -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datagen-reduce-dataset-snapshots, .gemini/skills/datagen-reduce-dataset-snapshots, .github/skills/datagen-reduce-dataset-snapshots and .opencode/skills/datagen-reduce-dataset-snapshots in your project.

What does Datagen Reduce Dataset Snapshots need to run?

Going by SKILL.md and its folder, Datagen Reduce Dataset Snapshots needs the command-line tools its instructions call (git, python and pip). Our summary lists: Python 3.

Does Datagen Reduce Dataset Snapshots access the network?

SKILL.md contains no URLs. Its commands use git and pip, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Datagen Reduce Dataset Snapshots safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Datagen Reduce Dataset Snapshots use?

Datagen Reduce Dataset Snapshots is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Datagen Reduce Dataset Snapshots use?

About 2.7k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Datagen Reduce Dataset Snapshots?

Skills that share tags, products or a category with Datagen Reduce Dataset Snapshots: Hoi Object Reconstruction Doctor (nvidia-isaac/video_to_data, 861 stars), CI CD And Automation (dzhalaevd/Donatello, 135 stars), Crabbox Quickstart (openclaw/crabbox, 1.5k stars) and Ego Reconstruction Setup (nvidia-isaac/video_to_data, 861 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Datagen Reduce Dataset Snapshots?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.