Agent skill

Datagen Standard Launch

by open-thoughts in open-thoughts/OpenThoughts-Agent

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Datagen Standard Launch

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-standard-launch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/datagen-standard-launch .claude/skills/datagen-standard-launch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
datagen-standard-launch
GitHub stars
301
Token cost
~947 tokens
SKILL.md length
209 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

  • Bulk completion/synthetic-data generation
  • SKILL.md covers Path 1 — Curator sharded…, Path 2 — declarative /… and Cleanup / verification
  • Calls python and conda
  • Tasks that involve LLM inference and serving

What it does

Datagen Standard Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. Two paths: (1) Curator sharded datagen — the multi-node data-parallel runcuratordatagensharded.sbatch (one vLLM server per node, disjoint dataset slices, auto-resume, manual afterany restart chain, curator conda env, --account=reformo); (2) the declarative generate.py / class-based generateabstract.py generators under data/ (data/generation BaseDataGenerator + InferenceEngine for…

Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving, Test data and fixtures and Autonomous loops. It works with vLLM, OpenAI and MiniMax. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Bulk completion/synthetic-data generation
  • Tasks that involve LLM inference and serving
  • Tasks that involve Test data and fixtures

Example prompts

  • “/datagen-standard-launch”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • conda

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Datagen Standard Launch loads about 947 tokens when it runs. Until then it costs about 188 tokens; SKILL.md has 209 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~188
When it runs · the whole SKILL.md, loaded when a task matches
~947

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 209 words, ~947 tokens.

Download SKILL.mdSave it as .claude/skills/datagen-standard-launch/SKILL.md (or your agent's skills folder).
name
datagen-standard-launch
description
Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. Two paths: (1) Curator sharded datagen — the multi-node data-parallel run_curator_datagen_sharded.sbatch (one vLLM server per node, disjoint dataset slices, auto-resume, manual afterany restart chain, `curator` conda env, --account=reformo); (2) the declarative `generate.py` / class-based `generate_abstract.py` generators under data/ (data/generation BaseDataGenerator + InferenceEngine for OpenAI/Anthropic/vLLM). Use for bulk completion/synthetic-data generation. The AGENTIC trace-generation path (Harbor + Daytona rollouts, MiniMax/GLM trace sets) is the SEPARATE `datagen-launch` skill.

datagen-standard-launch

Generate completions or synthetic data directly from a model, without Harbor or Daytona. For agentic trace-gen, use datagen-launch.


Path 1 — Curator sharded datagen (multi-node DP) — the primary path

data/sbatches/run_curator_datagen_sharded.sbatch runs one vLLM server per node over disjoint input slices and merges the results. Default: 32 nodes (#SBATCH --nodes=32, 4 GPUs/node), --account=reformo on Jupiter, and the curator environment (not otagent).

Auto-resume: stable shard output paths data/sbatches/curator_runs/<model>__<dataset>__<N>shards/shard_<i>/checkpoint_*.parquet let an afterany restart resume after a SLURM timeout.

Pre-req (login node, has internet): cache the input dataset first —

bash
conda activate curator
python -c "from datasets import load_dataset; ds=load_dataset('<dataset>',split='train'); print(len(ds))"

Launch — positional args <model> <input_dataset> <output_repo> [limit] [save_every]:

bash
# Simple (no restarts):
sbatch data/sbatches/run_curator_datagen_sharded.sbatch <model> <input_dataset> <output_repo> [limit] [save_every]

# With restart chain (recommended for long datasets) — build the afterany chain MANUALLY:
FIRST=$(sbatch data/sbatches/run_curator_datagen_sharded.sbatch \
  <model> <input_dataset> <output_repo> [limit] [save_every] | awk '{print $4}')
PREV=$FIRST; for i in $(seq 1 6); do
  PREV=$(sbatch --dependency=afterany:$PREV \
    data/sbatches/run_curator_datagen_sharded.sbatch \
    <model> <input_dataset> <output_repo> [limit] [save_every] | awk '{print $4}')
done
  • MAX_RESTARTS is not implemented; build the --dependency=afterany: chain by hand.
  • save_every: pass 700, not the default 200 (the fifth positional arg).
  • limit (4th arg) caps rows for a smoke run; omit for the full set.

Path 2 — declarative / class-based generator scripts (data/)

data/ has named pipeline directories in two styles:

  • Declarative generate.py — self-contained scripts for local / one-off runs:
    bash
    python data/<dataset>/generate.py [--flags]
  • Class-based generate_abstract.py — subclass BaseDataGenerator for HPC runs with launcher-managed vLLM endpoints, submitted through the unified launcher:
    bash
    python -m hpc.launch --job_type datagen \
      --datagen_script data/<dataset>/generate_abstract.py \
      --datagen_target_repo <org/repo> \
      --datagen_extra_args "--stage both --limit 2000"

Core modules in data/generation/: base.py (BaseDataGenerator), schemas.py (GenerationRequest/GenerationResult), engines.py (InferenceEngine for OpenAI / Anthropic / vLLM).


Cleanup / verification

Standard datagen pushes merged parquet shards to <output_repo>. Verify the row count and that the repo self-populated; no Daytona/trace-export step applies. Cluster details: .agents/ops/jupiter/; launcher: .agents/projects/ot-agent/ot-agent.md.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/datagen-standard-launch of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Datagen Standard Launch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Datagen Standard Launch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Datagen Standard Launch this skillopen-thoughts/OpenThoughts-Agent301—~947Automated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k3 repos~3kAutomated safety check: PassMIT
Model Serving MinefieldBlackwellboy/model-serving-minefield135—~2.1kAutomated safety check: PassMIT
vLLM Model ServingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Tanstack AIsecondsky/claude-skills2271 repos~3.6kAutomated safety check: NotesMIT
Vllm Bench Random Syntheticvllm-project/vllm-skills103—~1.5kAutomated safety check: PassApache-2.0

Similar skills

  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 3 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Model Serving Minefield

    Blackwellboy/model-serving-minefield

    Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.

    135 GitHub stars~2.1k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • vLLM Model Serving

    Orchestra-Research/AI-Research-SKILLs

    Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Tanstack AI

    secondsky/claude-skills

    TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.

    227 GitHub starsUsed in 1 repo~3.6k tokens
    AI & LLM EngineeringAuto-check: notes
  • Vllm Bench Random Synthetic

    vllm-project/vllm-skills

    Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.

    103 GitHub stars~1.5k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed
  • Vllm Bench Serve

    vllm-project/vllm-skills

    Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.

    103 GitHub stars~1.6k tokensUpdated 6 mo ago
    AI & LLM EngineeringAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 10 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 10 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 10 days ago
    Auto-check passed

Questions about Datagen Standard Launch

What does Datagen Standard Launch do?

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. Datagen Standard Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

When should I use Datagen Standard Launch?

Datagen Standard Launch fits situations like: bulk completion/synthetic-data generation; tasks that involve LLM inference and serving; tasks that involve Test data and fixtures.

How do I install Datagen Standard Launch in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a claude-code`. Or copy the skill folder (.agents/skills/datagen-standard-launch in open-thoughts/OpenThoughts-Agent) into .claude/skills/datagen-standard-launch in your project. Claude Code loads it when a task matches its description.

How do I install Datagen Standard Launch in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a codex`. Or copy the skill folder (.agents/skills/datagen-standard-launch in open-thoughts/OpenThoughts-Agent) into .agents/skills/datagen-standard-launch in your project. Codex loads it when a task matches its description.

Can I use Datagen Standard Launch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datagen-standard-launch, .gemini/skills/datagen-standard-launch, .github/skills/datagen-standard-launch and .opencode/skills/datagen-standard-launch in your project.

What does Datagen Standard Launch need to run?

Going by SKILL.md and its folder, Datagen Standard Launch needs the command-line tools its instructions call (python and conda). Our summary lists: Python 3.

Does Datagen Standard Launch access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Datagen Standard Launch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Datagen Standard Launch use?

Datagen Standard Launch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Datagen Standard Launch use?

About 947 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Datagen Standard Launch?

Skills that share tags, products or a category with Datagen Standard Launch: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Tanstack AI (secondsky/claude-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Datagen Standard Launch?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.