Agent skill

Crud Archive Run

by open-thoughts in open-thoughts/OpenThoughts-Agent

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

Apache-2.0Auto-check passedAI & LLM Engineering

Install Crud Archive Run

skills CLI
$ npx skills add open-thoughts/OpenThoughts-Agent --skill crud-archive-run -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install open-thoughts/OpenThoughts-Agent crud-archive-run --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/open-thoughts/OpenThoughts-Agent.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/crud-archive-run .claude/skills/crud-archive-run && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
crud-archive-run
GitHub stars
301
Token cost
~1.2k tokens
SKILL.md length
388 words
Files
1
Skills in repo
44
Repo updated
First seen
Licence
Apache-2.0

At a glance

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…

  • Concluding/archiving an experiment
  • SKILL.md covers ✅ ARCHIVE (always — every run…, ❌ SKIP (non-informative AND…, Mechanic — tar… and Per-run-type components — WHAT…, plus 1 more section
  • Calls rsync and aws
  • Before a cleanup skill rms an on-disk tree

What it does

Crud Archive Run is an agent skill from open-thoughts/OpenThoughts-Agent. Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL stdout/stderr (incl. vLLM/serving logs), and wandb. Pack-rat by design: if it's potentially informative, keep it. Only skip the non-informative-or-huge (model weights/checkpoints, core/memory dumps, massive raw tmux-pane / terminal-recording bytes). Many-tiny-files → tar THEN rsync (never rsync thousands of small files…

Its SKILL.md is about 1.2k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering LLM inference and serving. It works with vLLM, Weights & Biases and tmux. The repository describes itself as: Data recipes and robust infrastructure for training AI agents. The licence is Apache-2.0.

When your agent uses it

  • Concluding/archiving an experiment
  • Before a cleanup skill rms an on-disk tree
  • Before CoreWeave R2/pod artifacts get GCd

Example prompts

  • “/crud-archive-run”

What it can do on your machine

Read from SKILL.md and the folder at commit 3bd1917. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • rsync
    • aws

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use rsync and aws, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Crud Archive Run loads about 1.2k tokens when it runs. Until then it costs about 213 tokens; SKILL.md has 388 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~213
When it runs · the whole SKILL.md, loaded when a task matches
~1.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from open-thoughts/OpenThoughts-Agent at commit 3bd1917, republished under its Apache-2.0 licence (© open-thoughts). 388 words, ~1,245 tokens.

Download SKILL.mdSave it as .claude/skills/crud-archive-run/SKILL.md (or your agent's skills folder).
name
crud-archive-run
description
Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor trace_jobs (raw per-trial traces), ALL ray logs, ALL stdout/stderr (incl. vLLM/serving logs), and wandb. Pack-rat by design: if it's potentially informative, keep it. Only skip the non-informative-or-huge (model weights/checkpoints, core/memory dumps, massive raw tmux-pane / terminal-recording bytes). Many-tiny-files → tar THEN rsync (never rsync thousands of small files raw). Use when concluding/archiving an experiment, before a cleanup skill `rm`s an on-disk tree, or before CoreWeave R2/pod artifacts get GC'd. Per-run-type component maps live below; WHERE each artifact lives per cluster is a pointer into `.agents/ops/<cluster>/` and `.agents/projects/{harbor,marinskyrl,ot-agent}/`.

crud-archive-run

Archive potentially informative artifacts; skip only files that are both non-informative and large.

✅ ARCHIVE (always — every run type)

  • All raw Harbor traces — every per-trial dir under trace_jobs/ (RL) / eval_jobs/<name>/ (eval) / the datagen trace dir: result.json, config.json, manifest.json, lock.json, agent/trajectory.json (the raw agent transcript — INFORMATIVE, keep even at multi-MB), verifier/ (reward.txt, ctrf.json, test-stdout.txt), step_results, trajectory.summarization-*.json.
  • All ray logs — the ray session dir (raylet, gcs_server, per-worker *.out/*.err, python-core-*).
  • All stdout / stderr — SLURM .out/.err, the complete CoreWeave finelog, vllm.log, job.log, and per-trial trial.log.
  • wandb — the local wandb/ run dir if present; else record the run URL/id in the archive's MANIFEST.
  • Configs / launch command / rendered YAML / metric CSVs / trainer_log.jsonl.

❌ SKIP (non-informative AND large)

  • Model weights / checkpoints — *.safetensors, *.pt, *.bin, global_step_*/, consolidated shards (they live on HF / R2; not useful for post-hoc debugging).
  • Core / memory dumps — core.*, *.hprof, coredump trees.
  • Massive raw terminal-pane bytes — *.pane (raw tmux pane dumps) and agent/recording.cast (asciinema) only when large and redundant with trajectory.json; keep small casts.
  • Conda/uv/pip caches, extracted wheel trees, __pycache__, .venv.

Mechanic — tar many-small-files, then rsync

On the source cluster/pod, tar the small-file tree before rsyncing it:

# on the cluster (SLURM) — one tarball per run, excluding the SKIP set
tar --exclude='*.safetensors' --exclude='*.pt' --exclude='*.bin' --exclude='global_step_*' \
    --exclude='core.*' --exclude='*.pane' \
    -czf /tmp/<run>_archive.tgz -C <run_dir> trace_jobs logs *.log config* wandb  # adjust to what exists
rsync -aP <cluster>:/tmp/<run>_archive.tgz  <dest>/         # then rm the /tmp tarball

Keep large single logs (vllm.log) in the tarball or rsync them alongside. Verify the tarball is non-empty and lists the expected trees (tar tzf … | head) before deleting the source.

Show full SKILL.md (178 more words)Show less

Per-run-type components — WHAT + WHERE (pointers, they drift — read the ops/projects doc)

  • CoreWeave agentic RL (SkyRL/MarinSkyRL) — durable traces: s3://marin-us-east-02a/iris/<job>/trace_jobs (--trials-dir auto; pull with aws s3 --endpoint-url <R2>); pod-local traces: /app/experiments/<run>/trace_jobs (grab before pod GC with scripts/iris/analyze_coreweave_rl_job_live.sh <pod> cp). Full log: iris … job logs --since-ms <submit> --no-tail. Ray logs and vllm.log are pod-local; record the W&B URL.
  • SFT (LLaMA-Factory / axolotl, SLURM) — .out per-step logs at experiments/<job>/logs/*.out, trainer_log.jsonl, rendered config, wandb. Weights → HF (SKIP). Log path via scontrol show job <id> -o StdOut=/%Z. Details: .agents/projects/{llama-factory,axolotl}/, the cluster ops doc.
  • Datagen (Harbor traces) — one-level trace_jobs/<trial>/result.json + the harbor run log; the artifact is the trace set (→ HF), but archive the trace_jobs + logs. Details: .agents/projects/harbor/.
  • Eval (agentic Harbor) — eval_jobs/<name>/<trial>/{result.json,config.json,agent/trajectory.json, verifier/,manifest.json,lock.json} + top-level vllm.log, job.log, per-trial trial.log. Skip the big recording.cast/*.pane when large. Details: .agents/projects/harbor/, eval-agentic-cleanup.
  • Cluster paths: .agents/ops/iris/ (CoreWeave), .agents/ops/tacc/ (SLURM), and .agents/ops/empireai/ (SLURM). Read the relevant one first.

Destination

Default: ~/Documents/experiments/<active|complete>/<exp>/run_archive/<run-id>/. Write a one-line MANIFEST (run-id, cluster, job-id, dates, kept/skipped artifacts, W&B URL). Optionally push the tarball to HF (penfever/…-archive, public default laion/) or R2. Archive and verify before any cleanup reclaims the tree.

© open-thoughts, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/crud-archive-run of open-thoughts/OpenThoughts-Agent.

Open the folder on GitHubat commit 3bd1917

Compare with similar skills

Crud Archive Run next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Crud Archive Run compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Crud Archive Run this skillopen-thoughts/OpenThoughts-Agent301—~1.2kAutomated safety check: PassApache-2.0
SageMaker Serving Image Selectionhuggingface/skills11k1 repos~4.6kAutomated safety check: PassApache-2.0
Aider DelegateamElnagdy/delegate-skills2.3k2 repos~3kAutomated safety check: PassMIT
Hugging Face Local Model Evalshuggingface/skills11k2 repos~1.6kAutomated safety check: PassApache-2.0
Diffusion Perf Optvllm-project/vllm-omni7.1k—~7.5kAutomated safety check: PassApache-2.0
Ascend Release Manager for vLLMvllm-project/vllm-ascend2.9k—~7.2kAutomated safety check: PassApache-2.0

Similar skills

  • Official

    Chooses the right serving container and current image URI for deploying a Hugging Face model to a SageMaker endpoint, preferring Hugging Face images over generic ones.

    11k GitHub starsUsed in 1 repo~4.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Aider Delegate

    amElnagdy/delegate-skills

    Delegate a coding task to Aider (aider) as a background implementer, then review its diff and land it yourself.

    2.3k GitHub starsUsed in 2 repos~3k tokens
    AI & LLM EngineeringAuto-check passed
  • Official

    Runs evaluations of Hugging Face Hub models on local hardware with inspect-ai or lighteval, and helps choose between vLLM, Transformers and accelerate backends.

    11k GitHub starsUsed in 2 repos~1.6k tokens
    AI & LLM EngineeringAuto-check passed
  • Diffusion Perf Opt

    vllm-project/vllm-omni

    Diagnose and optimize vLLM Omni diffusion workloads, especially Wan/Qwen/Flux-style image and video generation.

    7.1k GitHub stars~7.5k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Ascend Release Manager for vLLM

    vllm-project/vllm-ascend

    Runs the end-to-end vLLM Ascend release process: opens the release checklist and feedback issues, scans for release-blocking bugs and test coverage gaps, and generates release notes and announcements.

    2.9k GitHub stars~7.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Quantization

    vllm-project/vllm-omni

    Work on vLLM-Omni quantization for diffusion, autoregressive, omni, or multi-stage models.

    7.1k GitHub stars~1.4k tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from open-thoughts/OpenThoughts-Agent

All 44 skills in this repo
  • Analyze Dataset Token Length

    open-thoughts/OpenThoughts-Agent

    Analyze the token length of an OT-Agent conversation-format (ShareGPT-style) dataset — the per-trace distribution (median/p90/max) and/or counts under a token threshold + a metadata predicate (e.g.

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Id Eval Ranking

    open-thoughts/OpenThoughts-Agent

    Given a list of models (HF name stubs) that have valid agentic ID eval scores in Supabase, build a ranking table: raw per-benchmark accuracy on the 3 ID benchmarks (SWE-Bench-100…

    301 GitHub stars~3.1k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Job History Iris

    open-thoughts/OpenThoughts-Agent

    Run the Iris harbor job-history analyzer (scripts/iris/analyzeirisharborjob.py) on a datagen/eval job and read its JSON sidecar for trustworthy throughput / preemption / productive-trial stats.

    301 GitHub stars~2.9k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Rl Behavior

    open-thoughts/OpenThoughts-Agent

    Run the full RL behavioral-analysis pipeline (scripts/analysis/analyzerlbehavior.py) on a trained RL model to understand WHAT changed vs its pre-RL baseline, WHY, whether it PERSISTS, and its EVAL…

    301 GitHub stars~4.2k tokensUpdated 11 days ago
    Auto-check passed
  • Analyze Training Run Iris

    open-thoughts/OpenThoughts-Agent

    Detailed health check for a Levanter/executor TRAINING run on the marin Iris cluster (e.g.

    301 GitHub stars~2k tokensUpdated 11 days ago
    Auto-check passed
  • Code Create Staged Plan

    open-thoughts/OpenThoughts-Agent

    DESIGN a non-trivial codebase change (Harbor / MarinSkyRL / vLLM / OT-Agent / LLaMA-Factory) as a dependency-ordered STAGED PLAN before writing code — a feature port, a multi-step fix with parity…

    301 GitHub stars~1.5k tokensUpdated 11 days ago
    Auto-check passed

Questions about Crud Archive Run

What does Crud Archive Run do?

Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL…. Crud Archive Run is an agent skill from open-thoughts/OpenThoughts-Agent. Durably ARCHIVE everything informative from a finished run / experiment before it's cleaned up or its cluster artifacts age out — ALL Harbor tracejobs (raw per-trial traces), ALL ray logs, ALL stdout/stderr (incl.

When should I use Crud Archive Run?

Crud Archive Run fits situations like: concluding/archiving an experiment; before a cleanup skill rms an on-disk tree; before CoreWeave R2/pod artifacts get GCd.

How do I install Crud Archive Run in Claude Code?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill crud-archive-run -a claude-code`. Or copy the skill folder (.agents/skills/crud-archive-run in open-thoughts/OpenThoughts-Agent) into .claude/skills/crud-archive-run in your project. Claude Code loads it when a task matches its description.

How do I install Crud Archive Run in Codex?

Run `npx skills add open-thoughts/OpenThoughts-Agent --skill crud-archive-run -a codex`. Or copy the skill folder (.agents/skills/crud-archive-run in open-thoughts/OpenThoughts-Agent) into .agents/skills/crud-archive-run in your project. Codex loads it when a task matches its description.

Can I use Crud Archive Run in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add open-thoughts/OpenThoughts-Agent --skill crud-archive-run -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/crud-archive-run, .gemini/skills/crud-archive-run, .github/skills/crud-archive-run and .opencode/skills/crud-archive-run in your project.

What does Crud Archive Run need to run?

Going by SKILL.md and its folder, Crud Archive Run needs the command-line tools its instructions call (rsync and aws).

Does Crud Archive Run access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Crud Archive Run safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Crud Archive Run use?

Crud Archive Run is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Crud Archive Run use?

About 1.2k tokens (SKILL.md is roughly 5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Crud Archive Run?

Skills that share tags, products or a category with Crud Archive Run: SageMaker Serving Image Selection (huggingface/skills, 11k stars), Aider Delegate (amElnagdy/delegate-skills, 2.3k stars), Hugging Face Local Model Evals (huggingface/skills, 11k stars) and Diffusion Perf Opt (vllm-project/vllm-omni, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Crud Archive Run?

open-thoughts (a GitHub organization) maintains it in open-thoughts/OpenThoughts-Agent, which has 301 GitHub stars. The repository holds 44 skills in this directory. The repository was last updated on September 28, 2026.

Source: open-thoughts/OpenThoughts-Agent on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.