Search
AI & LLM Engineering · open-thoughts/OpenThoughts-Agent
Skills
Sort:BestMost starsTrending todayTrending this weekTrending this monthNewestRecently updatedName
| # | Skill | Repository | Stars | Used in | Tokens | Auto-check | Licence | Updated |
|---|---|---|---|---|---|---|---|---|
| 1 | Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g. | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 2 | Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100… | open-thoughts/ | 301 | — | ~3.1k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 3 | DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity… | open-thoughts/ | 301 | — | ~1.5k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 4 | Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL… | open-thoughts/ | 301 | — | ~1.2k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 5 | Guardrailed DELETE of auto-registered eval sandboxjobs rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error 10%). | open-thoughts/ | 301 | — | ~2.8k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 6 | Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow… | open-thoughts/ | 301 | — | ~2.5k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 7 | Launch, monitor, and manually clean up a trajectory-generation (datagen) job on Marin's Iris TPU cluster via the OpenThoughts-Agent entrypoint. | open-thoughts/ | 301 | — | ~4.3k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 8 | Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes. | open-thoughts/ | 301 | — | ~947 | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 9 | Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy… | open-thoughts/ | 301 | — | ~3.7k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 10 | Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint. | open-thoughts/ | 301 | — | ~3.8k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 11 | Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints. | open-thoughts/ | 301 | — | ~2.3k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 12 | Preserve + publish a finished RL (SkyRL/GRPO) training checkpoint after the job terminates (completed at maxsteps OR early-stopped/scancelled) on an HPC cluster (Jupiter/Leonardo/Perlmutter). | open-thoughts/ | 301 | — | ~2.4k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 13 | Launch, relaunch, or sweep STANDARD (non-agentic) SkyRL RL on CINECA Leonardo — GRPO on math/reasoning datasets (gsm8k, MATH/aime) and on-policy distillation (OPD, teacher→student) — via raw sbatch… | open-thoughts/ | 301 | — | ~3.2k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 14 | Publish + clean up a finished LLaMA-Factory SFT job on a no-internet HPC cluster (Jupiter/Leonardo): cancel pending retries, drop intermediate checkpoints, HF-upload the model to its configured… | open-thoughts/ | 301 | — | ~1.9k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 15 | 15.Sft Launch Launch SFT via python -m hpc.launch --jobtype sft on any cluster (JSC Jupiter GH200, CINECA Leonardo A100, TACC Vista GH200), with EITHER backend — LLaMA-Factory (default) or axolotl (--sftbackend… | open-thoughts/ | 301 | — | ~2.9k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 16 | Read, aggregate, and (carefully) write OT-Agent eval/model data in the Supabase registry. | open-thoughts/ | 301 | — | ~6.2k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 17 | Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent. | open-thoughts/ | 301 | — | ~5.8k | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 18 | EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated… | open-thoughts/ | 301 | — | ~977 | Automated safety check: Pass | Apache-2.0 | 12 days ago |
| 19 | Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API. | open-thoughts/ | 301 | — | ~1.8k | Automated safety check: Warn | Apache-2.0 | 12 days ago |