Aider Delegate
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-standard-launch --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/datagen-standard-launch .claude/skills/datagen-standard-launch && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "datagen-standard-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-standard-launch into .claude/skills/datagen-standard-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-standard-launch", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-standard-launchType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-standard-launch --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/datagen-standard-launch .agents/skills/datagen-standard-launch && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "datagen-standard-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-standard-launch into .agents/skills/datagen-standard-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-standard-launch", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-standard-launch --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/datagen-standard-launch .cursor/skills/datagen-standard-launch && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "datagen-standard-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-standard-launch into .cursor/skills/datagen-standard-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-standard-launch", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/datagen-standard-launch--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-standard-launch --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/datagen-standard-launch .gemini/skills/datagen-standard-launch && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "datagen-standard-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-standard-launch into .gemini/skills/datagen-standard-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-standard-launch", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-standard-launchInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/datagen-standard-launch .github/skills/datagen-standard-launch && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "datagen-standard-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-standard-launch into .github/skills/datagen-standard-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-standard-launch", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-standard-launch --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/datagen-standard-launch .opencode/skills/datagen-standard-launch && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "datagen-standard-launch" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-standard-launch into .opencode/skills/datagen-standard-launch/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-standard-launch", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
datagen-standard-launchLaunch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.
Datagen Standard Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. Two paths: (1) Curator sharded datagen — the multi-node data-parallel runcuratordatagensharded.sbatch (one vLLM server per node, disjoint dataset slices, auto-resume, manual afterany restart chain, curator conda env, --account=reformo); (2) the declarative generate.py / class-based generateabstract.py generators under data/ (data/generation BaseDataGenerator + InferenceEngine for…
Its SKILL.md is about 950 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in AI & LLM Engineering, covering LLM inference and serving, Test data and fixtures and Autonomous loops. It works with vLLM, OpenAI and MiniMax. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythoncondaFrom the folder's file list and the shell code blocks in SKILL.md.
No URLs in SKILL.md.
From URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Datagen Standard Launch loads about 947 tokens when it runs. Until then it costs about 188 tokens; SKILL.md has 209 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 209 words, ~947 tokens.
.claude/skills/datagen-standard-launch/SKILL.md (or your agent's skills folder).Generate completions or synthetic data directly from a model, without Harbor or Daytona. For agentic trace-gen,
use datagen-launch.
data/sbatches/run_curator_datagen_sharded.sbatch runs one vLLM server per node over disjoint input slices and
merges the results. Default: 32 nodes (#SBATCH --nodes=32, 4 GPUs/node), --account=reformo on Jupiter,
and the curator environment (not otagent).
Auto-resume: stable shard output paths
data/sbatches/curator_runs/<model>__<dataset>__<N>shards/shard_<i>/checkpoint_*.parquet let an afterany
restart resume after a SLURM timeout.
Pre-req (login node, has internet): cache the input dataset first —
conda activate curator
python -c "from datasets import load_dataset; ds=load_dataset('<dataset>',split='train'); print(len(ds))"Launch — positional args <model> <input_dataset> <output_repo> [limit] [save_every]:
# Simple (no restarts):
sbatch data/sbatches/run_curator_datagen_sharded.sbatch <model> <input_dataset> <output_repo> [limit] [save_every]
# With restart chain (recommended for long datasets) — build the afterany chain MANUALLY:
FIRST=$(sbatch data/sbatches/run_curator_datagen_sharded.sbatch \
<model> <input_dataset> <output_repo> [limit] [save_every] | awk '{print $4}')
PREV=$FIRST; for i in $(seq 1 6); do
PREV=$(sbatch --dependency=afterany:$PREV \
data/sbatches/run_curator_datagen_sharded.sbatch \
<model> <input_dataset> <output_repo> [limit] [save_every] | awk '{print $4}')
doneMAX_RESTARTS is not implemented; build the --dependency=afterany: chain by hand.save_every: pass 700, not the default 200 (the fifth positional arg).limit (4th arg) caps rows for a smoke run; omit for the full set.data/)data/ has named pipeline directories in two styles:
generate.py — self-contained scripts for local / one-off runs:python data/<dataset>/generate.py [--flags]generate_abstract.py — subclass BaseDataGenerator for HPC runs with launcher-managed vLLM
endpoints, submitted through the unified launcher:python -m hpc.launch --job_type datagen \
--datagen_script data/<dataset>/generate_abstract.py \
--datagen_target_repo <org/repo> \
--datagen_extra_args "--stage both --limit 2000"Core modules in data/generation/: base.py (BaseDataGenerator), schemas.py
(GenerationRequest/GenerationResult), engines.py (InferenceEngine for OpenAI / Anthropic / vLLM).
Standard datagen pushes merged parquet shards to <output_repo>. Verify the row count and that the repo
self-populated; no Daytona/trace-export step applies. Cluster details: .agents/ops/jupiter/; launcher:
.agents/projects/ot-agent/ot-agent.md.
© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/datagen-standard-launch of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Datagen Standard Launch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Datagen Standard Launch this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~947 | Automated safety check: Pass | Apache-2.0 | |
| Aider DelegateamElnagdy/delegate-skills | 2.3k | 3 repos | ~3k | Automated safety check: Pass | MIT | |
| Model Serving MinefieldBlackwellboy/model-serving-minefield | 135 | — | ~2.1k | Automated safety check: Pass | MIT | |
| vLLM Model ServingOrchestra-Research/AI-Research-SKILLs | 13k | 6 repos | ~2.3k | Automated safety check: Pass | MIT | |
| Tanstack AIsecondsky/claude-skills | 227 | 1 repos | ~3.6k | Automated safety check: Notes | MIT | |
| Vllm Bench Random Syntheticvllm-project/vllm-skills | 103 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 |
amElnagdy/delegate-skills
Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.
Blackwellboy/model-serving-minefield
Diagnose OpenAI-compatible model-serving failures from symptoms, endpoint reports, explicit configuration files, or logs while preserving evidence status and requiring confirm/refute checks.
Orchestra-Research/AI-Research-SKILLs
Deploys LLMs with vLLM for high-throughput serving, covering the OpenAI-compatible server, offline batch inference, monitoring and a Docker rollout.
secondsky/claude-skills
TanStack AI (alpha) provider-agnostic type-safe chat with streaming for OpenAI, Anthropic, Gemini, Ollama.
vllm-project/vllm-skills
Run vLLM performance benchmark using synthetic random data to measure throughput, TTFT (Time to First Token), TPOT (Time per Output Token), and other key performance metrics.
vllm-project/vllm-skills
Benchmark vLLM or OpenAI-compatible serving endpoints using vllm bench serve.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
Categories
Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. Datagen Standard Launch is an agent skill from open-thoughts/OpenThoughts-Agent. Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.
Datagen Standard Launch fits situations like: bulk completion/synthetic-data generation; tasks that involve LLM inference and serving; tasks that involve Test data and fixtures.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a claude-code`. Or copy the skill folder (.agents/skills/datagen-standard-launch in open-thoughts/OpenThoughts-Agent) into .claude/skills/datagen-standard-launch in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a codex`. Or copy the skill folder (.agents/skills/datagen-standard-launch in open-thoughts/OpenThoughts-Agent) into .agents/skills/datagen-standard-launch in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-standard-launch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datagen-standard-launch, .gemini/skills/datagen-standard-launch, .github/skills/datagen-standard-launch and .opencode/skills/datagen-standard-launch in your project.
Going by SKILL.md and its folder, Datagen Standard Launch needs the command-line tools its instructions call (python and conda). Our summary lists: Python 3.
SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Datagen Standard Launch is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 947 tokens (SKILL.md is roughly 3.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Datagen Standard Launch: Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Model Serving Minefield (Blackwellboy/model-serving-minefield, 135 stars), vLLM Model Serving (Orchestra-Research/AI-Research-SKILLs, 13k stars) and Tanstack AI (secondsky/claude-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.