Stripe Projects
fossasia/eventyay
A skill your agent uses when the user wants to provision infrastructure or third-party services using Stripe Projects.
Post-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-job-cleanup --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/datagen-job-cleanup .claude/skills/datagen-job-cleanup && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "datagen-job-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-job-cleanup into .claude/skills/datagen-job-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-job-cleanup", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-job-cleanupType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-job-cleanup --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.agents/skills/datagen-job-cleanup .agents/skills/datagen-job-cleanup && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "datagen-job-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-job-cleanup into .agents/skills/datagen-job-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-job-cleanup", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-job-cleanup --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.agents/skills/datagen-job-cleanup .cursor/skills/datagen-job-cleanup && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "datagen-job-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-job-cleanup into .cursor/skills/datagen-job-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-job-cleanup", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/open-thoughts/OpenThoughts-Agent.git --path .agents/skills/datagen-job-cleanup--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-job-cleanup --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.agents/skills/datagen-job-cleanup .gemini/skills/datagen-job-cleanup && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "datagen-job-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-job-cleanup into .gemini/skills/datagen-job-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-job-cleanup", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-job-cleanupInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .github/skills && cp -r skills-src/.agents/skills/datagen-job-cleanup .github/skills/datagen-job-cleanup && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "datagen-job-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-job-cleanup into .github/skills/datagen-job-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-job-cleanup", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install open-thoughts/OpenThoughts-Agent datagen-job-cleanup --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.agents/skills/datagen-job-cleanup .opencode/skills/datagen-job-cleanup && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "datagen-job-cleanup" agent skill from https://github.com/open-thoughts/OpenThoughts-Agent/tree/main/.agents/skills/datagen-job-cleanup into .opencode/skills/datagen-job-cleanup/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "datagen-job-cleanup", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
datagen-job-cleanupPost-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk.
Datagen Job Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent. Post-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk. There is NO model checkpoint — the artifact is the trace dataset. Covers the TIMEOUT-strands-traces gotcha (uploads silently never run), the ONE-level tracejobs nesting (vs RL's double-nest), the real-vs-failed sanity check (avgturns ≈ 1.0 = dead run, don't upload), the otagent-env uploader, the non-empty HF verify, and…
Its SKILL.md is about 4.3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Backend & APIs, covering File uploads and storage. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Shell commands in SKILL.md call:
pythongsutilcurlpython3From the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
huggingface.coFrom URLs in SKILL.md, links to its own repository left out.
Names these keys or tokens, usually read from environment variables:
OPENAI_API_KEYHF_TOKENFrom names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Datagen Job Cleanup loads about 4.3k tokens when it runs. Until then it costs about 210 tokens; SKILL.md has 1,687 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 1,687 words, ~4,294 tokens.
.claude/skills/datagen-job-cleanup/SKILL.md (or your agent's skills folder).After a datagen (trace-generation) job terminates, follow these steps to get the traces onto HF and free disk. Unlike RL/SFT there is no model checkpoint — the artifact is the trace dataset.
Datagen jobs run to a wall-clock --time_limit; on TIMEOUT, Harbor's terminal trace-upload step is killed
mid-run, leaving completed trials on disk but no upload. A clean COMPLETED exit usually uploads
automatically; a TIMEOUT almost never does. Either way, do NOT assume the upload happened — verify (step 4).
sacct -j <jobid> -X --format=JobID,JobName%55,State,Elapsed,End -nDatagen writes to <run_dir>/trace_jobs/<inner_run_name>/<task>__<id>/ — a SINGLE level of nesting under
trace_jobs/, NOT the double-nest <run>/<run>/trace_jobs of the RL path. The --job_dir you pass to the
uploader must be the dir that DIRECTLY CONTAINS the <task>__<id> trial dirs, i.e. trace_jobs/<inner_run_name>:
# Direct-ssh launches land under experiments/; --experiments_dir launches under ot-baf.
RUN=/e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/<job_name> # or /e/data1/.../ot-baf/<job_name>
INNER=$(ls -d $RUN/trace_jobs/*/ 2>/dev/null | head -1); echo "$INNER"--s3-output-dir is a campaign/root prefix, not a per-job destination. The
shared Iris output resolver appends the Iris job name and routes Harbor to:
<s3-output-root>/<iris-job>/trace_jobs/<harbor-job>/<trial>/result.jsonFor example, launch with
--s3-output-dir s3://marin-us-east-02a/iris/<campaign>, not with a URI ending
in $JOB. Eval and datagen both use the IrisOutputPaths plan in
hpc/iris/outputs.py; do not reproduce this path construction in a launcher.
Run S3 cleanup inside a cw-rno2a Iris worker so the transfer uses cluster
credentials and bandwidth:
python scripts/harbor/cleanup_coreweave_datagen_s3.py \
--target '<job>|s3://marin-us-east-02a/iris/<campaign>/<job>|penfever/<repo>'The downloader lists the prefix once, downloads with bounded parallelism and
adaptive S3 retries, applies the realness gate, uploads in one Hub commit,
reloads the published train split, and removes only its worker-local temporary
copy. It never deletes the durable S3 source. Repeat --target to process
several completed jobs serially and idempotently.
When a dataset row has several independently launched -rN attempts, give
their durable trial_uri values as a comma-separated prefix list in one
target. The cleanup tool creates a separate local run directory for each
prefix, retains each run's config.json, then exports their trials into the
one requested Hub repository. Repeated task IDs are intentional independent
samples and are retained; only exact copied artifacts must not be counted
twice. Do not point the tool at only the newest retry.
python scripts/harbor/cleanup_coreweave_datagen_s3.py \
--target '<dataset>|s3://.../first-attempt,s3://.../retry-r2|penfever/<dataset>-traces'--no_literal_tokens, --single_commit, and --skip_register are cleanup
tool defaults; pass only --target arguments to this wrapper. Its reusable,
git-tracked implementation is
scripts/harbor/cleanup_coreweave_datagen_s3.py.
Jobs launched before the shared-path fix may have the malformed rescue layout
<root>/<iris-job>/<harbor-job>/<trial> with no trace_jobs component. The
cleanup tool recognizes this existing layout from result.json parents, but
new launchers must only write the canonical layout above.
A served /v1/models healthcheck does NOT mean generation worked. Compute avg turn count + exception
rate: if avg turns ≈ 1.0, the run is near-total failure (e.g. endpoint healthy + many result.json, but
every trial a 1-turn InternalServerError from a dead EngineCore). A real run has multi-step trajectories
(turns > 1) + a tolerable exception rate (tezos datagen ~20-25% AgentTimeout is normal).
# trajectory.json is a dict with a "steps" list; turns ≈ len(steps).
$OTAGENT_PY - <<'PY'
import json, glob, os, statistics
inner = os.environ["INNER"]
dirs = glob.glob(os.path.join(inner, "*__*/"))
turns, exc = [], 0
for d in dirs:
tj = os.path.join(d, "agent", "trajectory.json")
if os.path.exists(tj):
t = json.load(open(tj))
turns.append(len(t.get("steps", [])) if isinstance(t, dict) else len(t))
r = os.path.join(d, "result.json")
if os.path.exists(r):
if (json.load(open(r)).get("exception_info") or {}).get("exception_type"): exc += 1
print(f"trials={len(dirs)} avg_turns={statistics.mean(turns):.2f} exceptions={exc}" if turns else "no trajectories")
PYIf avg_turns ≈ 1.0 → the run failed; do NOT upload. Diagnose instead (read a trial's exception.txt + the
_vllm.log for the engine-side error) and write an agent_log; the traces are not worth keeping.
From the otagent conda env — the uploader needs google.cloud.storage + matplotlib, which envs/rl lacks:
# otagent env, source secrets first. $INNER = the dir DIRECTLY containing the <task>__<id> trial dirs.
python -m scripts.harbor.make_and_upload_trace_dataset \
--job_dir "$INNER" \
--repo_id penfever/<descriptive-name> \
--episodes last \
--served_model <the exact served model ref> \ # required for decodable literals
--include_literal_tokens \ # REQUIRE literals (fail loud if 0 bind)
--single_commit # ONE HF commit — avoids the 128-commit/hr 429 on big setsDefault to the penfever/ org + --episodes last. For a chunked launch, upload each chunk's trace_jobs to
its own _chunk{i} repo. Public by default (feedback_hf_public_default).
Literals are AUTO-INCLUDED when present (no flag needed). The uploader favors the durable
literal.jsonl: it auto-discovers the sibling <experiments_dir>/logs/*_literal.jsonl (searching
--job_dir and a few parents) and correlates the RecordProxy records into the trajectory step metrics so
the exported dataset carries the trainable prompt_token_ids / completion_token_ids / logprobs
columns. Requirements + knobs:
--job_dir must sit inside the experiments-dir tree so the parent-walk reaches …/logs/. $INNER
qualifies as long as the local run dir still has its sibling logs/ (local runs do).gs://marin-models-{us,eu}. Jobs pin to a
co-located single-region bucket (gs://marin-us-east5, …); older jobs are on the multi-region mirror.
Resolve each job's ACTUAL recorded OUTER prefix (registry-first, iris-fallback):JOB=<job_name>
OUT=$(/Users/benjaminfeuer/miniconda3/envs/otagent/bin/python -m hpc.iris.job_output_resolver "$JOB" \
--cluster /Users/benjaminfeuer/Documents/marin/lib/iris/config/marin.yaml) # e.g. gs://marin-us-east5/ot-agent/<job>
mkdir -p /tmp/${JOB}_traces
gsutil -m rsync -r "$OUT/" /tmp/${JOB}_traces/ # OUTER <job>/ — carries logs/*_literal.jsonl--cluster is only
consulted for the iris fallback when the job isn't in the local registry).<job>/ is the --job_dir, NOT the outer rescue root. You rsync the
OUTER $OUT/ → /tmp/<job>_traces so logs/*_literal.jsonl rides along, but the trial dirs sit one
level down at /tmp/<job>_traces/<job>/<trial>, so pass --job_dir **/tmp/<job>_traces/<job>** (the
inner subdir that DIRECTLY contains the <task>__<id> dirs). Passing the outer /tmp/<job>_traces still
uploads TEXT rows, but the extra <job>/ prefix makes 0 trials bind → silent text-only dataset (with
--include_literal_tokens it fails loud instead). Also: gsutil rsync needs the local dest to EXIST —
mkdir -p before the rsync.--single_commit (HF 128-commits/hour cap). ~1 commit per shard; a
multi-thousand-row set + any re-attempt blows past HF's 128 repo-commits/hour limit → 429 mid-push
→ a partial dataset. --single_commit pushes the whole dataset in ONE commit → never trips the limit.
(If you already 429'd, the repo may be partial/empty; wait ~1h for the window to reset, then re-run with
--single_commit.)messages from records after reconstruct_chains.literal.jsonl exports text-only, byte-identical to before (parity).--no_literal_tokens forces text-only even when a literal.jsonl is present.--include_literal_tokens means REQUIRE: fails loud if literals are expected but none are found.literal.jsonl is present but 0 trials bind (a regression, not a valid
dataset).--served_model on any literal upload. Token-id columns are only decodable with the EXACT
tokenizer the engine served (a same-family tokenizer decodes word tokens to garbage); the uploader stamps
the model ref into tokenizer_provenance.json + the dataset-card README. For the opencode-131k campaign:
--served_model Qwen/Qwen3.5-122B-A10B-FP8 (or gs://marin-models-us/ot-agent/models/Qwen/Qwen3.5-122B-A10B-FP8/
— storage root via marin_prefix(), see .agents/ops/iris/ops.md §rendezvous, don't
hardcode the region bucket).7c978b78): the exporter pins the literal token columns to an explicit nested
type per shard. Datasets uploaded BEFORE 7c978b78 are degraded (under-populated literal yield, shards
with no token columns, load_dataset CastError) — re-rescue them to recover full yield. Full
reference (decoding, tokenizer provenance, re-rescue, literals→SFT): .agents/projects/harbor/ops.md
(§ Literal-token trace datasets).If numeric-result counts, agent/trajectory.json counts, and SFT-exported row
counts differ materially, do not infer the cause from aggregate yield. Run the
git-tracked analyzer against every retry prefix or local Harbor job directory:
python -m scripts.harbor.analyze_sft_export_gaps \
--dataset s3://.../first-attempt \
--dataset s3://.../retry-r2 \
--output /tmp/<dataset>-sft-export-gap-analysis.json
# Local equivalent:
python -m scripts.harbor.analyze_sft_export_gaps \
--dataset "$INNER" \
--output /tmp/<dataset>-sft-export-gap-analysis.jsonFor S3 inputs, run inside a cw-rno2a Iris worker with the normal object-store
credentials so downloads remain in-region. The analyzer joins job-level reward
buckets to per-trial result.json and trajectory artifacts, preserves retry/run
provenance, and invokes the same conversation collection and dataset conversion
path as the exporter. Its JSON report accounts for both boundaries separately:
agent/trajectory.json;agent/trajectory.json → collected conversation → SFT-ready dataset row.Use the mutually exclusive reason counts and per-trial evidence to remediate
recoverable failures before uploading. Preserve the report in the experiment
artifacts or an agent_logs/ escalation when the gap cannot be repaired. Do not
declare the numeric gate cleared from an aggregate counter until this join is
consistent.
The repo may exist as a 0-row shell (a prior failed/partial upload, or Harbor pre-creating it); an existing repo is NOT proof of success. Confirm row count:
curl -s -H "Authorization: Bearer $HF_TOKEN" \
"https://huggingface.co/api/datasets/penfever/<descriptive-name>" \
| python3 -c "import sys,json; d=json.load(sys.stdin); print('files:', len(d.get('siblings',[])), 'lastMod:', d.get('lastModified'))"The uploader's own "Generating train split: N examples" line is ground truth — N must match the real (non-1-turn) trial count from step 2. Zero files / 0 rows = the upload did not land; re-run step 3.
⚠ Re-uploading over a repo that had a prior TEXT-ONLY upload. A leftover
train-00000-of-00001.parquetfrom the earlier text-only push has a no-token schema that can MASK the new token-bearing shards →count_populated_literal_rowsreads 0 even though the re-upload landed. If a post-upload verify shows 0 populated literals but the uploader logged a goodLiteral yield, delete the stale text-only parquet(s) (huggingface_hub.HfApi().delete_file) and re-verify. Cleaner: re-upload to a fresh repo id to avoid stale-shard masking entirely.
Also verify the literal columns landed (jobs run with --record_literal). The uploader prints a
[trace-export] Literal yield: X/Y trials … line — X should be > 0. Confirm in the pushed dataset:
import datasets
from scripts.harbor.make_and_upload_trace_dataset import count_populated_literal_rows
ds = datasets.load_dataset("penfever/<descriptive-name>", split="train")
print("rows w/ literals:", count_populated_literal_rows(ds.data.table)) # must be > 0 for a --record_literal job0 populated rows on a --record_literal job = the literals dropped — re-run step 3 with the correct
--job_dir (must reach the sibling logs/), or pass --literal_log <gs://…/logs/<slug>_literal.jsonl>.
rm -rf $RUN…/ot-baf/tmp_tasks/<name> made for a chunking smoke test, or a
scripts.datagen.extract_tasks_from_parquet output you won't reuse. Do NOT delete the shared canonical
task dirs under /e/data1/datasets/playground/ot/tasks/<benchmark> — reused across runs.rm -rf /e/data1/datasets/playground/ot-baf/tmp_tasks/<one-off-name> # only if one-offexperiments/active/datagen/minimax-m2.7-tt/tracker.md. Per-dataset cycle:
extract tasks → launch chunks (chunk_size 500, --chunk_array_max 5, single-node, verifier-ON) → wait
ALL chunks COMPLETE (validate via sacct, not empty squeue) → consolidate _chunk{N} via
join_hf_repos.py + verify row count == sum, each chunk near-full, no run_id gaps, realness
(avg_turns>1) → DELETE chunk repos (only after consolidated verified) → clean local dirs → launch next
row. Target repo: penfever/<source-basename>-minimax-m27-131k-traces. Only hard auto-skip =
oversized-extraction rows (≫10k rows busting the inode budget, e.g. knowledge-mcqa 616k) → ask.
Snapshots: do NOT reclaim per-dataset (shared envs); track cumulative tally vs the org cap (see daytona
doc). The 3h cron is the natural driver.openai/gpt-4o-mini) can be silently reward-dead (all 0.0) if
OPENAI_API_KEY is invalid/revoked → judge 401 → reward defaults 0.0 while traces are genuine
(avg_turns healthy). Tell at consolidation: verifier_output has judge API error … 401 on ~all rows.
Check the REWARD DISTRIBUTION at consolidation, not just avg_turns. Rows generated before the
2026-06-11 fix (e.g. #22) stay reward-dead unless re-scored. User precedent: accept reward-dead rows as
turns-only SFT data + advance (record the caveat in the tracker row). Test a key:
curl …/v1/chat/completions -H "Authorization: Bearer $OPENAI_API_KEY" … → expect 200.© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
Just SKILL.md in .agents/skills/datagen-job-cleanup of open-thoughts/OpenThoughts-Agent.
Open the folder on GitHubat commit 3bd1917
Datagen Job Cleanup next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Datagen Job Cleanup this skillopen-thoughts/OpenThoughts-Agent | 301 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | |
| Stripe Projectsfossasia/eventyay | 1.7k | 5 repos | ~2k | Automated safety check: Notes | Apache-2.0 | |
| FoundatioFoundatioFx/Foundatio | 2.1k | — | ~3.9k | Automated safety check: Pass | Apache-2.0 | |
| Spatialduckdb/duckdb-skills | 603 | 1 repos | ~1k | Automated safety check: Notes | MIT | |
| R Oopab604/claude-code-r-skills | 207 | 2 repos | ~1.9k | Automated safety check: Pass | MIT | |
| Edgestore Setupedgestorejs/edgestore | 454 | — | ~1.8k | Automated safety check: Pass | MIT |
fossasia/eventyay
A skill your agent uses when the user wants to provision infrastructure or third-party services using Stripe Projects.
FoundatioFx/Foundatio
A skill your agent uses when working with Foundatio infrastructure abstractions for .NET -- caching, queuing, messaging, file storage, distributed locking, or background jobs.
duckdb/duckdb-skills
Answer questions about spatial data using DuckDB. An agent skill from duckdb/duckdb-skills.
ab604/claude-code-r-skills
R object-oriented programming guide for S7, S3, S4, and vctrs.
edgestorejs/edgestore
A skill your agent uses when adding, extending, or troubleshooting EdgeStore file uploads, upload UI, or bucket access policies in a TypeScript/React application.
duckdb/duckdb-skills
Explore and query data on S3, Cloudflare R2, GCS, MinIO, or any S3-compatible storage.
open-thoughts/OpenThoughts-Agent
Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.
open-thoughts/OpenThoughts-Agent
Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…
open-thoughts/OpenThoughts-Agent
Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.
open-thoughts/OpenThoughts-Agent
Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…
open-thoughts/OpenThoughts-Agent
Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.
open-thoughts/OpenThoughts-Agent
DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…
Categories
Post-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk. Datagen Job Cleanup is an agent skill from open-thoughts/OpenThoughts-Agent. Post-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk.
Datagen Job Cleanup fits situations like: A datagen/trace job finishes (COMPLETED; TIMEOUT) and its traces need uploading + verifying; consolidating a chunked datagen run.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a claude-code`. Or copy the skill folder (.agents/skills/datagen-job-cleanup in open-thoughts/OpenThoughts-Agent) into .claude/skills/datagen-job-cleanup in your project. Claude Code loads it when a task matches its description.
Run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a codex`. Or copy the skill folder (.agents/skills/datagen-job-cleanup in open-thoughts/OpenThoughts-Agent) into .agents/skills/datagen-job-cleanup in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill datagen-job-cleanup -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/datagen-job-cleanup, .gemini/skills/datagen-job-cleanup, .github/skills/datagen-job-cleanup and .opencode/skills/datagen-job-cleanup in your project.
Going by SKILL.md and its folder, Datagen Job Cleanup needs the command-line tools its instructions call (python, gsutil, curl and python3) and credentials named OPENAI_API_KEY and HF_TOKEN. Our summary lists: Python 3.
SKILL.md names 1 domain. In commands or code: huggingface.co; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
Datagen Job Cleanup is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 4.3k tokens (SKILL.md is roughly 17k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Datagen Job Cleanup: Stripe Projects (fossasia/eventyay, 1.7k stars), Foundatio (FoundatioFx/Foundatio, 2.1k stars), Spatial (duckdb/duckdb-skills, 603 stars) and R Oop (ab604/claude-code-r-skills, 207 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.
Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.