Repository

open-thoughts/OpenThoughts-Agent agent skills

Every skill in the open-thoughts/OpenThoughts-Agent repository on GitHub, ranked by score, with the commands to install them.
skills
44
GitHub stars
301

GitHub description: “Data recipes and robust infrastructure for training AI agents”

Stars
301 (43 forks)
Licence
Apache-2.0
Last push
Sep 2026
Created
Dec 2025

Install all skills

skills CLI (any agent)
npx skills add open-thoughts/OpenThoughts-Agent

Add --skill <name> for a single skill and -a <agent> to choose the agent (see the agent guides).

Skills in open-thoughts/OpenThoughts-Agent, ranked

Ranked by score. Sort bymost stars,trending,newest,recently updated

Skills in open-thoughts/OpenThoughts-Agent, ranked
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

open-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.09 days ago
2

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

open-thoughts/OpenThoughts-Agent301—~3.1kAutomated safety check: PassApache-2.09 days ago
3

Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

open-thoughts/OpenThoughts-Agent301—~2.9kAutomated safety check: PassApache-2.09 days ago
4

Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

open-thoughts/OpenThoughts-Agent301—~4.2kAutomated safety check: PassApache-2.09 days ago
5

Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

open-thoughts/OpenThoughts-Agent301—~2kAutomated safety check: PassApache-2.09 days ago
6

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

open-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.09 days ago
7

Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format.

open-thoughts/OpenThoughts-Agent301—~2.2kAutomated safety check: NotesApache-2.09 days ago
8

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

open-thoughts/OpenThoughts-Agent301—~1.2kAutomated safety check: PassApache-2.09 days ago
9

Guardrailed DELETE of auto-registered eval sandboxjobs rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error 10%).

open-thoughts/OpenThoughts-Agent301—~2.8kAutomated safety check: PassApache-2.09 days ago
10

Safely purge stale, never-populated sandboxjobs placeholder rows (eval launches that died/stalled before scoring) from the OT-Agent Supabase registry.

open-thoughts/OpenThoughts-Agent301—~2.9kAutomated safety check: PassApache-2.09 days ago
11

Post-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk.

open-thoughts/OpenThoughts-Agent301—~4.3kAutomated safety check: PassApache-2.09 days ago
12

Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow…

open-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.09 days ago
13

Launch, monitor, and manually clean up a trajectory-generation (datagen) job on Marin's Iris TPU cluster via the OpenThoughts-Agent entrypoint.

open-thoughts/OpenThoughts-Agent301—~4.3kAutomated safety check: PassApache-2.09 days ago
14

Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality.

open-thoughts/OpenThoughts-Agent301—~2.7kAutomated safety check: PassApache-2.09 days ago
15

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

open-thoughts/OpenThoughts-Agent301—~947Automated safety check: PassApache-2.09 days ago
16

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.09 days ago
17

Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint.

open-thoughts/OpenThoughts-Agent301—~3.8kAutomated safety check: PassApache-2.09 days ago
18

Consolidate FINISHED standard / lmeval (evalchemy) math-suite eval jobs — the Delphi 6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker.

open-thoughts/OpenThoughts-Agent301—~1.7kAutomated safety check: PassApache-2.09 days ago
19

Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints.

open-thoughts/OpenThoughts-Agent301—~2.3kAutomated safety check: PassApache-2.09 days ago
20

Format HPC job-status reports as box-drawing tables, bucketed by job type (RL · SFT · Datagen · Eval · Catch-all), with the right metric columns, signal thresholds, and red-flags per bucket.

open-thoughts/OpenThoughts-Agent301—~4kAutomated safety check: PassApache-2.09 days ago
21

Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost.

open-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.09 days ago
22

Re-register the every-3-hours UNIFIED OPS TICK cron — the CURRENT operator-owned monitor for the qwen3.5-122b-131k-datagen-opencode campaign (keep-3 datagen with autonomous rescue+refill) AND the…

open-thoughts/OpenThoughts-Agent301—~2.6kAutomated safety check: PassApache-2.09 days ago
23

Preserve + publish a finished RL (SkyRL/GRPO) training checkpoint after the job terminates (completed at maxsteps OR early-stopped/scancelled) on an HPC cluster (Jupiter/Leonardo/Perlmutter).

open-thoughts/OpenThoughts-Agent301—~2.4kAutomated safety check: PassApache-2.09 days ago
24

Launch / relaunch agentic RL (SkyRL terminalbench + Harbor + Daytona) on JSC Jupiter (GH200).

open-thoughts/OpenThoughts-Agent301—~2.7kAutomated safety check: PassApache-2.09 days ago
25

Launch, relaunch, or sweep STANDARD (non-agentic) SkyRL RL on CINECA Leonardo — GRPO on math/reasoning datasets (gsm8k, MATH/aime) and on-policy distillation (OPD, teacher→student) — via raw sbatch…

open-thoughts/OpenThoughts-Agent301—~3.2kAutomated safety check: PassApache-2.09 days ago
26

Publish + clean up a finished LLaMA-Factory SFT job on a no-internet HPC cluster (Jupiter/Leonardo): cancel pending retries, drop intermediate checkpoints, HF-upload the model to its configured…

open-thoughts/OpenThoughts-Agent301—~1.9kAutomated safety check: PassApache-2.09 days ago
27

Launch SFT via python -m hpc.launch --jobtype sft on any cluster (JSC Jupiter GH200, CINECA Leonardo A100, TACC Vista GH200), with EITHER backend — LLaMA-Factory (default) or axolotl (--sftbackend…

open-thoughts/OpenThoughts-Agent301—~2.9kAutomated safety check: PassApache-2.09 days ago
28

Delete stale Daytona sandboxes across ALL THREE orgs (DataComp, DataCompData, DataCompRL) in one pass.

open-thoughts/OpenThoughts-Agent301—~819Automated safety check: PassApache-2.09 days ago
29

Fix file permissions on a directory tree on an HPC cluster (Leonardo, Jupiter, TACC, etc.) so other users can read your shared data, conda envs, or work directories.

open-thoughts/OpenThoughts-Agent301—~714Automated safety check: PassApache-2.09 days ago
30

Reclaim idle Daytona SNAPSHOTS org-wide to free space under the 60-snapshot cap, using scripts/daytona/daytonasnapshotmanager.py.

open-thoughts/OpenThoughts-Agent301—~969Automated safety check: PassApache-2.09 days ago
31

Read, aggregate, and (carefully) write OT-Agent eval/model data in the Supabase registry.

open-thoughts/OpenThoughts-Agent301—~6.2kAutomated safety check: PassApache-2.09 days ago
32

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

open-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.09 days ago
33

Build a clean per-dataset summary table/CSV for a datagen (trajectory-generation) campaign — one row per task source with Status (COMPLETED / FAILED / RUNNING / NOT STARTED), N Trials Completed…

open-thoughts/OpenThoughts-Agent301—~1.9kAutomated safety check: PassApache-2.09 days ago
34

EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated…

open-thoughts/OpenThoughts-Agent301—~977Automated safety check: PassApache-2.09 days ago
35

File a GitHub issue for a bug or improvement found this session.

open-thoughts/OpenThoughts-Agent301—~1.6kAutomated safety check: PassApache-2.09 days ago
36

Produce a comprehensive cross-cluster job-status update for a recurring N-hourly cluster sweep.

open-thoughts/OpenThoughts-Agent301—~4.2kAutomated safety check: PassApache-2.09 days ago
37

The PROCEDURE for one every-3-hours Iris job-status sweep — primarily the marin TPU datagen/eval jobs ("iris" here = the marin TPU cluster), plus CoreWeave GPU-RL as monitor-only.

open-thoughts/OpenThoughts-Agent301—~3.3kAutomated safety check: PassApache-2.09 days ago
38

Re-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost.

open-thoughts/OpenThoughts-Agent301—~4.1kAutomated safety check: PassApache-2.09 days ago
39

Clean up a completed NON-AGENTIC / HF-only SFT model — HF upload WITHOUT Supabase DB registration.

open-thoughts/OpenThoughts-Agent301—~1.1kAutomated safety check: PassApache-2.09 days ago
40

Write or revise tests with an emphasis on behavior, regression coverage, pytest style, and avoiding "slop tests." Use when adding tests, fixing failing tests, reviewing test quality, or deciding…

open-thoughts/OpenThoughts-Agent301—~498Automated safety check: PassApache-2.09 days ago
41

Marin house writing style. An agent skill from open-thoughts/OpenThoughts-Agent.

open-thoughts/OpenThoughts-Agent301—~1.3kAutomated safety check: PassApache-2.09 days ago
42

Deep single-RL-job health probe → a KILL / NO-KILL / ERROR recommendation for the supervisor.

open-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.09 days ago
43

Debug a code bug with a structured debug log that records hypotheses, changes, and results.

open-thoughts/OpenThoughts-Agent301—~352Automated safety check: PassApache-2.09 days ago
44

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

open-thoughts/OpenThoughts-Agent301—~1.8kAutomated safety check: WarnApache-2.09 days ago

Questions, answered from the data.

What is the best skill in open-thoughts/OpenThoughts-Agent?

Analyze Dataset Token Length from open-thoughts/OpenThoughts-Agent ranks first of the 44 skills in open-thoughts/OpenThoughts-Agent listed here, with the highest score: its repository has 301 GitHub stars, its SKILL.md loads about 1.5k tokens and it passes the automated safety check with no findings. Next come Analyze Id Eval Ranking and Analyze Job History Iris.

Are the skills in open-thoughts/OpenThoughts-Agent official?

None yet. All 44 skills in open-thoughts/OpenThoughts-Agent listed here come from community repositories; a skill counts as official when the product's own GitHub organization publishes it.

How do I install all skills from open-thoughts/OpenThoughts-Agent?

Run npx skills add open-thoughts/OpenThoughts-Agent in your project: the open-source skills CLI installs the repository's skills into your coding agent's skills folder. To install a single skill, open its page here for the exact command.

How are these skills ranked?

By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.