Agent skill

Launch

by fcakyon in fcakyon/phd-skills

Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

MITAuto-check passedAI & LLM Engineering

Install Launch

skills CLI
$ npx skills add fcakyon/phd-skills --skill launch -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install fcakyon/phd-skills launch --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/fcakyon/phd-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/plugin/skills/launch .claude/skills/launch && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
launch
GitHub stars
414
Token cost
~1.4k tokens
SKILL.md length
711 words
Files
1
Skills in repo
12
Repo updated
First seen
Licence
MIT

At a glance

Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

  • Works in 5 steps: Config diff against a reference run → Run name discipline → Path verification → …
  • The user asks to launch
  • SKILL.md covers When to run, The checklist, Restart and kill cleanup and Output
  • Calls python; needs WANDB_API_KEY and NEPTUNE_API_TOKEN

What it does

Launch is an agent skill from fcakyon/phd-skills. Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup. Use when the user asks to launch, kick off, start, restart, or kill a training run, or mentions launching a multi-hour or multi-day GPU job (python train, accelerate launch, torchrun, deepspeed, sbatch, tmux training).

Its SKILL.md is about 1.4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. It works with Python and tmux. The repository describes itself as: PhD Research Skills for Claude Code: paper reproduction, experiment design, paper review, result comparison and more. The licence is MIT.

When your agent uses it

  • The user asks to launch
  • Kill a training run
  • Mentions launching a multi-hour
  • Multi-day GPU job (python train

Example prompts

  • “/launch”

Requirements

  • Python 3
  • A credential in WANDB_API_KEY
  • A credential in NEPTUNE_API_TOKEN

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Config diff against a reference run
  2. Run name discipline
  3. Path verification
  4. Monitoring setup
  5. ETA in your local timezone

What it can do on your machine

Read from SKILL.md and the folder at commit 67acd61. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names these keys or tokens, usually read from environment variables:

    • WANDB_API_KEY
    • NEPTUNE_API_TOKEN

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Launch loads about 1.4k tokens when it runs. Until then it costs about 93 tokens; SKILL.md has 711 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~93
When it runs · the whole SKILL.md, loaded when a task matches
~1.4k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from fcakyon/phd-skills at commit 67acd61, republished under its MIT licence (© fcakyon). 711 words, ~1,445 tokens.

Download SKILL.mdSave it as .claude/skills/launch/SKILL.md (or your agent's skills folder).
name
launch
description
Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup. Use when the user asks to launch, kick off, start, restart, or kill a training run, or mentions launching a multi-hour or multi-day GPU job (python train, accelerate launch, torchrun, deepspeed, sbatch, tmux training).

Launch: pre-flight checklist for long ML training jobs

Long training jobs are expensive to fail. A 12-hour run that crashes on epoch 3 from a missing dataset path or a default workers=8 against an NFS mount is a full day lost. This skill walks five quick checks before you commit the GPUs.

The agentic Stop hook in this plugin will route here from reason when an assistant tries to launch a run without going through the checklist.

When to run

The user just asked to:

  • launch / kick off / start / fire up a training run
  • restart a run that died
  • kill a current run (also runs the cleanup half of the checklist)
  • review a launch command before submitting

Or the user is about to run any of: python train.py, accelerate launch, torchrun, deepspeed, sbatch train.sh, tmux new-session ... python ... train, wandb sweep.

The checklist

1. Config diff against a reference run

The most expensive failure is launching with the wrong knobs. Before starting:

bash
find configs/ recipes/ experiments/ -maxdepth 3 \( -name '*.yaml' -o -name '*.yml' -o -name '*.json' -o -name '*.toml' \) -mtime -30 2> /dev/null | head

Pick the most-recently-modified config that resembles the intended run (same model family, same task). Diff against the intended config:

bash
diff -u configs/baseline_v1.yaml configs/intended.yaml

Walk every diff line. For each, ask: is this difference intentional and motivated, or is it a stale default I forgot to set? Common silent regressors:

  • num_workers / dataloader workers (default in many repos is 8: wrong on NFS)
  • batch_size (per-device vs global mismatch under DDP)
  • learning_rate (linearly scaled with batch size; if batch changed, lr should too)
  • optimizer betas / weight decay (paper-default vs framework-default)
  • mixed_precision (fp16 vs bf16 matters for some models)
  • gradient_accumulation_steps
  • seed (still set if you care about reproducibility)

If no reference exists in this project, ask the user to point at one. Do not launch with framework defaults alone.

2. Run name discipline

The run name will live in wandb / neptune / checkpoint dirs / status reports for the rest of its life. It must describe the experiment in plain English without internal codes:

  • bad: run-1, wave-2, cs-ad, phase2-internal
  • good: 7src-fastvit-s-featmap-mlp-dinov3, coco-baseline-bs256-lr3e-4, swin-t-imagenet-distill-from-vit-l

The pattern: <dataset/task>-<model>-<key-config>-<distinctive-recipe-piece>. If you can't describe the experiment from the name in one sentence, the name is wrong. The Stop hook flags any run reference that uses session-local labels.

3. Path verification

Before launching, every path the run depends on must be confirmed to exist:

bash
# Dataset path
ls -la /path/to/dataset | head

# Pretrained checkpoint (if loading)
ls -la /path/to/checkpoint.pt

# Output directory parent (must exist; the run dir will be created)
ls -la /path/to/runs/

# Config file
cat configs/intended.yaml | head

Never trust a path that was recalled from memory. The destructive_path_guard.sh hook will already block obvious cases for rm/mv, but the launch path needs the same scrutiny, a run started with a nonexistent dataset path crashes 30 minutes in instead of immediately.

Show full SKILL.md (299 more words)Show less
4. Monitoring setup

Auto-detect the experiment tracker:

  • WANDB_API_KEY set or wandb import in the launcher → wandb
  • NEPTUNE_API_TOKEN set → neptune
  • MLFLOW_TRACKING_URI set or mlflow in launcher → mlflow
  • presence of runs/ or lightning_logs/ → tensorboard
  • none of the above → ask the user; "no monitoring" is rarely the right answer for a multi-hour run

Confirm the run will appear under the right project / entity / experiment-name. Confirm any tags / groups for cohort comparison are set.

5. ETA in your local timezone

Estimate wall-clock duration: epochs × seconds-per-epoch / 3600 = hours. State the ETA in your local TZ (the system's TZ, which the timezone_scrub.sh hook validates against). If the run will straddle a meeting / sleep / OOO window, decide whether to defer or split.

Restart and kill cleanup

If this is a restart of a previously-failed run, or a kill before launching a replacement, purge stale artifacts in this exact order:

  1. Local checkpoint dir on the launching machine: rm -rf /local/runs/<run-name> (verify path first; the destructive_path_guard.sh will warn).
  2. Remote artifact dir on the cluster / NFS / object store: rm -rf /remote/runs/<run-name> (or equivalent).
  3. Experiment tracker run: delete via the tracker's API (wandb api.run(...).delete(), neptune run.stop() + delete via UI, etc.). Stale tracker runs corrupt later comparisons.
  4. Scheduler reservation: cancel the SLURM job (scancel <jobid>), the lambda labs reservation, the cron entry, etc. Runs that "killed but the GPUs are still allocated" are a recurring waste.

Skipping any of these creates ghost state that will confuse the next launch or the next comparison.

Output

When the user invokes this skill, walk the five checks (or three checks + cleanup, if killing) and report which passed and which failed. Block the launch on any failure unless the user explicitly waives the check.

For a clean launch, end with the launch command itself in a fenced block, ready to copy.

© fcakyon, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in plugin/skills/launch of fcakyon/phd-skills.

Open the folder on GitHubat commit 67acd61

Compare with similar skills

Launch next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Launch compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Launch this skillfcakyon/phd-skills414—~1.4kAutomated safety check: PassMIT
Paddle BuildPaddlePaddle/Paddle24k—~1kAutomated safety check: PassApache-2.0
Paddle Design CompilerPaddlePaddle/Paddle24k—~3.6kAutomated safety check: PassApache-2.0
Onnxtxtonnx/onnx22k—~1.3kAutomated safety check: PassApache-2.0
Benchmark Pyreflyfacebook/pyrefly7.1k—~1.8kAutomated safety check: PassMIT
Document Public APIspytorch/pytorch104k—~4.2kAutomated safety check: PassCustom licence

Similar skills

  • Paddle Build

    PaddlePaddle/Paddle

    A skill your agent uses when needing to compile, rebuild, or install Paddle from source after code changes.

    24k GitHub stars~1k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Compiler

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle 3.0 compiler full pipeline: SOT (Symbolic Opcode Translator) for bytecode-level dy2st graph capture, PIR (Paddle IR) for SSA-based intermediate…

    24k GitHub stars~3.6k tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed
  • Onnxtxt

    onnx/onnx

    Read or write ONNX text format ("onnxtxt"). An agent skill from onnx/onnx.

    22k GitHub stars~1.3k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Benchmark Pyrefly

    facebook/pyrefly

    Official

    Run Pyrefly benchmarks locally via Buck or Cargo, including PyTorch real-world LSP benchmarks.

    7.1k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Document Public APIs

    pytorch/pytorch

    Document undocumented public APIs in PyTorch by removing functions from coverageignorefunctions and coverageignoreclasses in docs/source/conf.py, running Sphinx coverage, and adding the appropriate…

    104k GitHub stars~4.2k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • ExecuTorch Cortex-M Backend

    pytorch/executorch

    Developer guide for the Cortex-M (CMSIS-NN) backend in ExecuTorch: quantization pipeline, pass manager, tests and adding new ops.

    5.1k GitHub stars~872 tokensUpdated today
    AI & LLM EngineeringAuto-check passed

More from fcakyon/phd-skills

All 12 skills in this repo
  • Reproduce

    fcakyon/phd-skills

    End-to-end paper reproduction from arxiv URL through smoke runs to replication experiments.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check passed
  • Compare

    fcakyon/phd-skills

    Same-epoch comparison of training runs across wandb, neptune, tensorboard, or mlflow.

    414 GitHub stars~1.2k tokensUpdated 21 days ago
    Auto-check passed
  • Debug

    fcakyon/phd-skills

    Evidence-before-action diagnosis of failing ML experiments. An agent skill from fcakyon/phd-skills.

    414 GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed
  • Experiment Design

    fcakyon/phd-skills

    A skill your agent uses when the user wants to design experiments, plan ablation studies, structure baselines, or create incremental evaluation strategies.

    414 GitHub stars~987 tokensUpdated 21 days ago
    Auto-check passed
  • Latex Setup

    fcakyon/phd-skills

    A skill your agent uses when the user wants to set up or troubleshoot a LaTeX environment, choose between biber and bibtex, install packages for a specific venue template, or configure compilation.

    414 GitHub stars~1.1k tokensUpdated 21 days ago
    Auto-check: notes
  • Literature Research

    fcakyon/phd-skills

    A skill your agent uses when the user wants to find related work, survey a research area, identify literature gaps, or discover open-source implementations.

    414 GitHub stars~999 tokensUpdated 21 days ago
    Auto-check passed

Works with

Questions about Launch

What does Launch do?

Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup. Launch is an agent skill from fcakyon/phd-skills. Pre-flight checklist for long-running ML training jobs covering config diff, run naming, path verification, monitoring setup, and restart-cleanup.

When should I use Launch?

Launch fits situations like: the user asks to launch; kill a training run; mentions launching a multi-hour; multi-day GPU job (python train.

How do I install Launch in Claude Code?

Run `npx skills add fcakyon/phd-skills --skill launch -a claude-code`. Or copy the skill folder (plugin/skills/launch in fcakyon/phd-skills) into .claude/skills/launch in your project. Claude Code loads it when a task matches its description.

How do I install Launch in Codex?

Run `npx skills add fcakyon/phd-skills --skill launch -a codex`. Or copy the skill folder (plugin/skills/launch in fcakyon/phd-skills) into .agents/skills/launch in your project. Codex loads it when a task matches its description.

Can I use Launch in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add fcakyon/phd-skills --skill launch -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/launch, .gemini/skills/launch, .github/skills/launch and .opencode/skills/launch in your project.

What does Launch need to run?

Going by SKILL.md and its folder, Launch needs the command-line tools its instructions call (python) and credentials named WANDB_API_KEY and NEPTUNE_API_TOKEN. Our summary lists: Python 3; A credential in WANDB_API_KEY; A credential in NEPTUNE_API_TOKEN.

Does Launch access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Launch safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Launch use?

Launch is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Launch use?

About 1.4k tokens (SKILL.md is roughly 5.8k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Launch?

Skills that share tags, products or a category with Launch: Paddle Build (PaddlePaddle/Paddle, 24k stars), Paddle Design Compiler (PaddlePaddle/Paddle, 24k stars), Onnxtxt (onnx/onnx, 22k stars) and Benchmark Pyrefly (facebook/pyrefly, 7.1k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Launch?

fcakyon (a GitHub user) maintains it in fcakyon/phd-skills, which has 414 GitHub stars. The repository holds 12 skills in this directory. The repository was last updated on September 16, 2026.

Source: fcakyon/phd-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.