Search

AI & LLM Engineering · By open-thoughts

19 skills found.
Search results
#SkillRepositoryStarsUsed inTokensAuto-checkLicenceUpdated
1

Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

open-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.012 days ago
2

Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

open-thoughts/OpenThoughts-Agent301—~3.1kAutomated safety check: PassApache-2.012 days ago
3

DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

open-thoughts/OpenThoughts-Agent301—~1.5kAutomated safety check: PassApache-2.012 days ago
4

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

open-thoughts/OpenThoughts-Agent301—~1.2kAutomated safety check: PassApache-2.012 days ago
5

Guardrailed DELETE of auto-registered eval sandboxjobs rows that DID score but FAILED the harvest gate — partial evals (valid-complete <90% or non-benign infra-error 10%).

open-thoughts/OpenThoughts-Agent301—~2.8kAutomated safety check: PassApache-2.012 days ago
6

Launch a datagen (trace-generation) job on an HPC cluster (Jupiter/Leonardo/Perlmutter) via the OpenThoughts-Agent hpc.launch --jobtype datagen entrypoint — the cluster-AGNOSTIC general flow…

open-thoughts/OpenThoughts-Agent301—~2.5kAutomated safety check: PassApache-2.012 days ago
7

Launch, monitor, and manually clean up a trajectory-generation (datagen) job on Marin's Iris TPU cluster via the OpenThoughts-Agent entrypoint.

open-thoughts/OpenThoughts-Agent301—~4.3kAutomated safety check: PassApache-2.012 days ago
8

Launch NON-AGENTIC (standard) data generation — plain vLLM/API completion generation with NO Harbor agent loop or Daytona sandboxes.

open-thoughts/OpenThoughts-Agent301—~947Automated safety check: PassApache-2.012 days ago
9

Launch agentic Harbor evals through the OT-Agent unified eval listener (eval/unifiedevallistener.py) on any cluster: select models (queryunevaledmodels.py / priority lists), wire the pinggy…

open-thoughts/OpenThoughts-Agent301—~3.7kAutomated safety check: PassApache-2.012 days ago
10

Launch, monitor, and manually clean up an eval job on Marin's Iris TPU or CoreWeave H100x8 GPU cluster via the OpenThoughts-Agent entrypoint.

open-thoughts/OpenThoughts-Agent301—~3.8kAutomated safety check: PassApache-2.012 days ago
11

Launch the fixed Delphi 6279 RL-scaling-laws downstream MATH eval suite (MATH-500 / AIME24 / gsm8k via evalchemy + lmeval) on CINECA Leonardo, for completed SFT / RL / base checkpoints.

open-thoughts/OpenThoughts-Agent301—~2.3kAutomated safety check: PassApache-2.012 days ago
12

Preserve + publish a finished RL (SkyRL/GRPO) training checkpoint after the job terminates (completed at maxsteps OR early-stopped/scancelled) on an HPC cluster (Jupiter/Leonardo/Perlmutter).

open-thoughts/OpenThoughts-Agent301—~2.4kAutomated safety check: PassApache-2.012 days ago
13

Launch, relaunch, or sweep STANDARD (non-agentic) SkyRL RL on CINECA Leonardo — GRPO on math/reasoning datasets (gsm8k, MATH/aime) and on-policy distillation (OPD, teacher→student) — via raw sbatch…

open-thoughts/OpenThoughts-Agent301—~3.2kAutomated safety check: PassApache-2.012 days ago
14

Publish + clean up a finished LLaMA-Factory SFT job on a no-internet HPC cluster (Jupiter/Leonardo): cancel pending retries, drop intermediate checkpoints, HF-upload the model to its configured…

open-thoughts/OpenThoughts-Agent301—~1.9kAutomated safety check: PassApache-2.012 days ago
15

Launch SFT via python -m hpc.launch --jobtype sft on any cluster (JSC Jupiter GH200, CINECA Leonardo A100, TACC Vista GH200), with EITHER backend — LLaMA-Factory (default) or axolotl (--sftbackend…

open-thoughts/OpenThoughts-Agent301—~2.9kAutomated safety check: PassApache-2.012 days ago
16

Read, aggregate, and (carefully) write OT-Agent eval/model data in the Supabase registry.

open-thoughts/OpenThoughts-Agent301—~6.2kAutomated safety check: PassApache-2.012 days ago
17

Audit + recover a finished agentic eval. An agent skill from open-thoughts/OpenThoughts-Agent.

open-thoughts/OpenThoughts-Agent301—~5.8kAutomated safety check: PassApache-2.012 days ago
18

EXECUTE a staged codebase plan (from code-create-staged-plan or an existing notes/<codebase/ plan) one stage at a time, gate-by-gate, while keeping the local clone ground truth and a dated…

open-thoughts/OpenThoughts-Agent301—~977Automated safety check: PassApache-2.012 days ago
19

Stand up a PUBLIC, shareable inference endpoint for an HF/gs model on an Iris TPU so people on the internet can vibe-test it in a browser or via an OpenAI-compatible API.

open-thoughts/OpenThoughts-Agent301—~1.8kAutomated safety check: WarnApache-2.012 days ago