GitHub organization
Agent skills by open-thoughts
- skills
- 44
- repository
- 1
Repositories by open-thoughts
Skills by open-thoughts, ranked
Ranked by score. Sort bymost stars,trending,newest,recently updated
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 2 | Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100… | open-thoughts/ | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 3 | Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats. | open-thoughts/ | 301 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 4 | Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL… | open-thoughts/ | 301 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 5 | Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g. | open-thoughts/ | 301 | — | ~2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 6 | DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity… | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 7 | 7.Commit Lint, run the pre-PR checks, commit, push, and author or update the branch's pull request in the required plain-text format. | open-thoughts/ | 301 | — | ~2.2k | Automated safety check: Notes | Apache-2.0 | 9 days ago |
| 8 | Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL… | open-thoughts/ | 301 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 9 | Guardrailed DELETE of auto-registered eval sandboxjobs rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error 10%). | open-thoughts/ | 301 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 10 | Safely purge stale, never-populated sandboxjobs placeholder rows (eval launches that died/stalled before scoring) from the OT-Agent Supabase registry. | open-thoughts/ | 301 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 11 | Post-run cleanup for a datagen (trace-generation) job on Iris/CoreWeave or an HPC cluster (Jupiter/Leonardo/Perlmutter): get the generated traces onto HF (penfever org) and free temporary disk. | open-thoughts/ | 301 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 12 | Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow… | open-thoughts/ | 301 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 13 | Launch, monitor, and manually clean up a trajectory-generation (datagen) job on Marin's Iris TPU cluster via the OpenThoughts-Agent entrypoint. | open-thoughts/ | 301 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 14 | Reduce the Daytona snapshot (unique-environment) count of a Harbor task dataset below the cap by editing its patcher's environment-build logic, without breaking task quality. | open-thoughts/ | 301 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 15 | Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. | open-thoughts/ | 301 | — | ~947 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 16 | Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy… | open-thoughts/ | 301 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 17 | Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint. | open-thoughts/ | 301 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 18 | Consolidate FINISHED standard / lmeval (evalchemy) math-suite eval jobs — the Delphi 6279 MATH-500 / AIME24 / gsm8k grid launched via eval-standard-launch — into the SCORES.md tracker. | open-thoughts/ | 301 | — | ~1.7k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 19 | Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints. | open-thoughts/ | 301 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 20 | Format HPC job-status reports as box-drawing tables, bucketed by job type (RL · SFT · Datagen · Eval · Catch-all), with the right metric columns, signal thresholds, and red-flags per bucket. | open-thoughts/ | 301 | — | ~4k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 21 | Re-register the every-3-hours Iris job-monitor cron (status check + datagen auto-rescue/keep-2-in-flight) if it has been lost. | open-thoughts/ | 301 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 22 | Re-register the every-3-hours UNIFIED OPS TICK cron — the CURRENT operator-owned monitor for the qwen3.5-122b-131k-datagen-opencode campaign (keep-3 datagen with autonomous rescue+refill) AND the… | open-thoughts/ | 301 | — | ~2.6k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 23 | Preserve + publish a finished RL (SkyRL/GRPO) training checkpoint after the job terminates (completed at maxsteps OR early-stopped/scancelled) on an HPC cluster (Jupiter/Leonardo/Perlmutter). | open-thoughts/ | 301 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 24 | Launch / relaunch agentic RL (SkyRL terminalbench + Harbor + Daytona) on JSC Jupiter (GH200). | open-thoughts/ | 301 | — | ~2.7k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 25 | Launch, relaunch, or sweep STANDARD (non-agentic) SkyRL RL on CINECA Leonardo — GRPO on math/reasoning datasets (gsm8k, MATH/aime) and on-policy distillation (OPD, teacher→student) — via raw sbatch… | open-thoughts/ | 301 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 26 | Publish + clean up a finished LLaMA-Factory SFT job on a no-internet HPC cluster (Jupiter/Leonardo): cancel pending retries, drop intermediate checkpoints, HF-upload the model to its configured… | open-thoughts/ | 301 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 27 | 27.Sft Launch Launch SFT via python -m hpc.launch --jobtype sft on any cluster (JSC Jupiter GH200, CINECA Leonardo A100, TACC Vista GH200), with EITHER backend — LLaMA-Factory (default) or axolotl (--sftbackend… | open-thoughts/ | 301 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 28 | Delete stale Daytona sandboxes across ALL THREE orgs (DataComp, DataCompData, DataCompRL) in one pass. | open-thoughts/ | 301 | — | ~819 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 29 | Fix file permissions on a directory tree on an HPC cluster (Leonardo, Jupiter, TACC, etc.) so other users can read your shared data, conda envs, or work directories. | open-thoughts/ | 301 | — | ~714 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 30 | Reclaim idle Daytona SNAPSHOTS org-wide to free space under the 60-snapshot cap, using scripts/daytona/daytonasnapshotmanager.py. | open-thoughts/ | 301 | — | ~969 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 31 | Read, aggregate, and (carefully) write OT-Agent eval/model data in the Supabase registry. | open-thoughts/ | 301 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 32 | Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent. | open-thoughts/ | 301 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 33 | Build a clean per-dataset summary table/CSV for a datagen (trajectory-generation) campaign — one row per task source with Status (COMPLETED / FAILED / RUNNING / NOT STARTED), N Trials Completed… | open-thoughts/ | 301 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 34 | EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated… | open-thoughts/ | 301 | — | ~977 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 35 | 35.File Issue File a GitHub issue for a bug or improvement found this session. | open-thoughts/ | 301 | — | ~1.6k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 36 | Produce a comprehensive cross-cluster job-status update for a recurring N-hourly cluster sweep. | open-thoughts/ | 301 | — | ~4.2k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 37 | The PROCEDURE for one every-3-hours Iris job-status sweep — primarily the marin TPU datagen/eval jobs ("iris" here = the marin TPU cluster), plus CoreWeave GPU-RL as monitor-only. | open-thoughts/ | 301 | — | ~3.3k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 38 | Re-create the local 3-hour tri-cluster cluster-sweep loop (Leonardo + CoreWeave(iris) + TACC(Vista); Jupiter SKIPPED until ~Jul 12) — the autonomous ML-ops monitor — if it has been lost. | open-thoughts/ | 301 | — | ~4.1k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 39 | Clean up a completed NON-AGENTIC / HF-only SFT model — HF upload WITHOUT Supabase DB registration. | open-thoughts/ | 301 | — | ~1.1k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 40 | 40.Write Tests Write or revise tests with an emphasis on behavior, regression coverage, pytest style, and avoiding "slop tests." Use when adding tests, fixing failing tests, reviewing test quality, or deciding… | open-thoughts/ | 301 | — | ~498 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 41 | Marin house writing style. An agent skill from open-thoughts/OpenThoughts-Agent. | open-thoughts/ | 301 | — | ~1.3k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 42 | Deep single-RL-job health probe → a KILL / NO-KILL / ERROR recommendation for the supervisor. | open-thoughts/ | 301 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 43 | 43.Debug Debug a code bug with a structured debug log that records hypotheses, changes, and results. | open-thoughts/ | 301 | — | ~352 | Automated safety check: Pass | Apache-2.0 | 9 days ago |
| 44 | Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. | open-thoughts/ | 301 | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | 9 days ago |
Questions, answered from the data.
What is the best skill by open-thoughts?
Analyze Dataset Token Length from open-thoughts/OpenThoughts-Agent ranks first of the 44 skills by open-thoughts listed here, with the highest score: its repository has 301 GitHub stars, its SKILL.md loads about 1.5k tokens and it passes the automated safety check with no findings. Next come Analyze Id Eval Ranking and Analyze Job History Iris.
Are open-thoughts's skills official?
None yet. All 44 skills by open-thoughts listed here come from community repositories; a skill counts as official when the product's own GitHub organization publishes it.
How are these skills ranked?
By Skill Navigator score, which combines the GitHub stars of the skill's repository (shared across that repo's skills and discounted for large collections), how many other GitHub owners carry a copy of the skill, and automated SKILL.md quality checks, minus penalties for safety-check warnings and for each further skill from the same repository. Skills that fail the safety check are not listed.