Agent skill

Derecho Jcm Runs

by climate-analytics-lab in climate-analytics-lab/jax-gcm

Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Derecho Jcm Runs

skills CLI
$ npx skills add climate-analytics-lab/jax-gcm --skill derecho-jcm-runs -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install climate-analytics-lab/jax-gcm derecho-jcm-runs --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/derecho-jcm-runs .claude/skills/derecho-jcm-runs && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
derecho-jcm-runs
GitHub stars
108
Token cost
~2.5k tokens
SKILL.md length
1,223 words
Files
7 (incl. scripts)
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues.

  • Works in 9 steps: Generate and submit → Always pre-flight before burning a queue… → Environment (baked into generated scripts) → …
  • Running any jcm model integration
  • SKILL.md covers 1. Generate and submit, 2. Always pre-flight before…, 3. Environment (baked into… and 4. Input data, plus 5 more sections
  • Runs Python and Shell scripts from its folder; calls python

What it does

Derecho Jcm Runs is an agent skill from climate-analytics-lab/jax-gcm. Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues. Use when running any jcm model integration, timestep/resolution sweep, performance benchmark, or GPU job on Derecho — covers job-script generation, queue/account selection, environment setup, and reliable completion monitoring.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 8 other files, including scripts (for example `reference/data_paths.md`, `scripts/mkjob.py` and `scripts/mkjob_test.py`).

It sits in AI & LLM Engineering, covering Deep learning. The repository describes itself as: GCM Physics written in JAX. The licence is Apache-2.0.

When your agent uses it

  • Running any jcm model integration
  • Timestep/resolution sweep
  • Performance benchmark
  • GPU job on Derecho — covers job-script generation

Example prompts

  • “/derecho-jcm-runs”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

9 steps, taken from the step headings in SKILL.md.

  1. Generate and submit
  2. Always pre-flight before burning a queue slot
  3. Environment (baked into generated scripts)
  4. Input data
  5. PBS facts specific to this machine
  6. Monitoring (scripts/watch_job.sh)
  7. Reading throughput correctly
  8. Benchmarking or debugging a performance difference
  9. Memory guidance (A100-40GB)

What it can do on your machine

Read from SKILL.md and the folder at commit 0940e89. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships 5 files in scripts/ (Python and Shell), which the agent can run.

    Shell commands in SKILL.md call:

    • python

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Derecho Jcm Runs loads about 2.5k tokens when it runs. Until then it costs about 83 tokens; SKILL.md has 1,223 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~83
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.

SKILL.md

The full file from climate-analytics-lab/jax-gcm at commit 0940e89, republished under its Apache-2.0 licence (© climate-analytics-lab). 1,223 words, ~2,520 tokens.

Download SKILL.mdSave it as .claude/skills/derecho-jcm-runs/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
derecho-jcm-runs
description
Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues. Use when running any jcm model integration, timestep/resolution sweep, performance benchmark, or GPU job on Derecho — covers job-script generation, queue/account selection, environment setup, and reliable completion monitoring.

Running jcm on Derecho

Layering. This is the Derecho/PBS site layer. The site-agnostic model layer (config groups, Hydra traps, stability overrides) is jcm-run, and throughput methodology is jcm-benchmark; both apply here too. The shared-workstation counterpart is devbox-jcm-runs — worth a glance for the contrast, since there GPUs are self-allocated rather than scheduled — and the Kubernetes counterpart is kubernetes-jcm-runs.

Derecho's A100s are 40 GB, half the cluster and dev-box cards. Throughput matches at equal work, but the memory ceiling does not: the reference table in jcm-benchmark is measured on 80 GB, and only its smallest configs fit here.

Generate a PBS script with scripts/mkjob.py, sanity-check the config, submit, then monitor with the patterns below. Every default here was established by a real campaign; the failure modes listed are ones that have actually happened.

1. Generate and submit

bash
python scripts/mkjob.py --name my_run --days 30 > runs/my_run.pbs
qsub runs/my_run.pbs

Common flags (see python scripts/mkjob.py --help for all):

flagdefaultnotes
--gpus N1>1 adds +grid.spmd_mesh and drops the memory fraction
--queuemainmain routes to gpu; gpudev for <1 h debugging
--hours6walltime
--gridecham_t63_l47_hybridany jcm/config/grid/*.yaml stem
--dt15minutes
--off-centering0.2SL off-centering; transport is always semi-Lagrangian
--physicsecham-jamecham-rrtmgp-2m for no aerosol
--radiation(config default)grey for a cheap-radiation A/B
--aquaplanetoffskips terrain/forcing files
--resumeoffreuse the run dir's checkpoint (without it the job deletes it)
--freshoffrefuse to generate if the run dir already has a checkpoint
--datamirrorHF bundles, prefetched at generation; local = legacy prepared files
--erapdpd (2005–2014) or pi (1850s) mirror climatologies
--emissions(local mode)legacy emissions file for --data local
--extra "k=v ..."—raw Hydra overrides appended last

2. Always pre-flight before burning a queue slot

Two checks, both cheap, both catch failures that otherwise waste a job:

bash
# (a) Hydra composition — catches +/++ prefix errors and unknown keys
JAX_PLATFORMS=cpu python -m jcm.main <exact overrides> --cfg job >/dev/null

# (b) coords constructibility — --cfg job does NOT build coords, so an
#     invalid spectral truncation only fails at runtime
JAX_PLATFORMS=cpu python -c "
from jcm.utils import get_coords
from jcm.physics.echam.echam_levels import get_echam_levels
get_coords(vertical_coords=get_echam_levels(<layers>), spectral_truncation=<T>)"

mkjob.py --check runs (a) for you and prints the command for (b).

Hydra override prefixes are a recurring trap: a key that already exists in the composed config takes no +; one that does not, requires it. run=longrun replaces the whole run group, so run.checkpoint_path needs + under it but not under the default run config.

3. Environment (baked into generated scripts)

bash
source ~/.venvs/jaxgcm/bin/activate
export PYTHONPATH=$REPO                   # jcm worktree wins over the venv's editable install
export JAX_PLATFORMS=cuda,cpu
export MAM4_JAX_ENABLE_X64=0              # f32 MAM4 core (forward-only); f64 default is much slower
export XLA_PYTHON_CLIENT_MEM_FRACTION=0.93   # 0.85 when ngpus>1 — 0.93 starves CUDA command buffers

Overridable site paths: JCM_REPO, JCM_VENV, JAM_INPUTS, JCM_EMISSIONS, PBS_ACCOUNT, SCRATCH.

Transport is always semi-Lagrangian (the Eulerian path was removed); the venv's dinosaur must be >= 1.5.0 (requirements.txt). Without it the dycore raises a clear install-instruction error; there is no fallback.

4. Input data

reference/data_paths.md lists every data source. The default is the HF data mirror (--data mirror --era pd|pi): mkjob.py derives every bundle path from --grid — terrain, forcing, emissions, DMS, dust, plus level-resolved ozone and oxidants from bundles/<grid>_l<levels>/ — and prefetches them on the login node at generation time, baking the local cache paths into the job. Compute nodes need no internet, and every grid/level combination the mirror carries (t63/t106/t127/t255 × l47/l95) works the same way: no packaged-grid special cases, no purge-eligible scratch files, and a grid/level mismatch fails at generation, not in the queue.

--data local keeps the legacy prepared-file behaviour (JAM_INPUTS / JCM_EMISSIONS, existence-checked before qsub). Its inputs are all grid-specific — level-resolved (ozone, oxidants) or horizontally validated (emissions, DMS, dust). forcing.ozone_file: auto resolves the packaged climatology (T63L47) and then the mirror's per-grid bundle, and on a hybrid grid raises if neither resolves rather than substituting the analytic profile (~7.6x the tropospheric ozone column). Prefer the mirror.

5. PBS facts specific to this machine

  • GPU account is UCSD0085 (UCSD0044 is casper-only and is rejected).
  • gpu_type=a100 must be inside the select chunk, not a separate -l.
  • -q main is a routing queue that lands GPU jobs in gpu; gpudev exists for short interactive-style debugging.
  • qsub -v VAR=x does not reach the job environment here — generated scripts hardcode their variables.
  • Keep #PBS -m abe so job mail keeps working (it was silently lost once when a script was derived by sed from one that omitted it).
  • Use set -euo pipefail; without -e a failed run still reaches a trailing touch DONE and looks successful.

6. Monitoring (scripts/watch_job.sh)

bash
scripts/watch_job.sh <jobid> <logfile> "<completion marker>"

Use it as the command of a persistent Monitor. It encodes five lessons:

  1. Read the log once per check. Grep the log into a variable, then both decide and report from those same bytes. Live NFS logs give stale re-reads, which produced repeated phantom "failures" whose detail printed empty.
  2. Debounce: a failure signature must persist across two checks.
  3. 3-strike qstat: PBS requeues and transient qstat errors otherwise look like a vanished job.
  4. File existence is not success: the driver writes chunk netCDFs before the NaN check. Verify the NaN vars: 0/N health line instead.
  5. Match verdicts, not keywords: jcm.main echoes the whole composed config on stdout, so every log contains bail_on_unhealthy: true. A bare unhealthy in FAIL_RE therefore failed every clean run.

Filter Lmod's "unknown module" noise — it is harmless on these nodes. scripts/watch_job_test.py (standalone, like mkjob_test.py — pytest does not collect dotted directories) checks both halves: a clean log passes and a real verdict still fails.

Show full SKILL.md (413 more words)Show less

7. Reading throughput correctly

Full methodology is in jcm-benchmark; the short version is that the N sim days/hr line in the log is cumulative and includes compile, so it must not be quoted. Use Wall: X s this chunk, discard chunk 1, and quote a rate only once the last two chunks agree.

bash
scripts/settled_rate.py <PBS stdout log> [--dt 15]   # per-chunk walls + convergence-checked rate

Give it the job's stdout log (runs/<tag>.log) — the Wall: X s this chunk lines are prints, so Hydra's main.log in the rundir has none of them. No PYTHONPATH is needed from either the in-repo or the installed ~/.claude/skills copy: the script finds the repo's tools/ by searching upward from itself and the working directory (JCM_REPO overrides).

That script and tools/benchmark.py share tools/chunk_timing.py, so the same run cannot yield two different answers. settled_rate.py reads a log that already exists (what you want for a job back from the queue); benchmark.py drives a run and samples GPU telemetry alongside (what you want on an interactive box).

Log locations differ by job type: a plain run writes to the PBS -o file (<name>.log in the submit directory); --bench variants write to $RUNDIR/<tag>/run.log. A 10-day run yields only two chunks and the analyzer will correctly refuse to quote a rate — allow >= 20 days (4 chunks) for a number worth reporting.

Reference points at T63L47, JAM + SL, dt=15, one A100-40GB: 151 s per 5 days = 119 days/hr, of which radiation is ~78%. Grey radiation gives ~34 s / 533 days/hr. See docs/source/design/dinosaur_sl_jam_configuration.md in the repo.

8. Benchmarking or debugging a performance difference

--bench emits a variant-matrix job (reference / grey radiation / any extra override sets) with convergence checks and GPU sampling under load. When comparing machines, capture on both: nvidia-smi static specs, clocks/power under load, Clocks Event Reasons, dependency provenance including git HEADs of editable installs, and the dtypes the model actually runs in — a config flag is not enough, since one f64 input promotes whole subgraphs. Power draw is diagnostic: high power at max clocks with low throughput indicates FP64 units engaging.

9. Memory guidance (A100-40GB)

T63L47 JAM fits comfortably at fraction 0.93 with 1 saved frame per chunk (2 frames OOM'd). T63L95 fits on one GPU. T106L95 does not — use 4 GPUs with +grid.spmd_mesh=[2,2,1] and fraction 0.85. Valid spectral truncations are 21, 31, 42, 63, 85, 106, 119, 127, 170, 213, 255, 340, 425. T127/T255 are ECHAM's own grids — supported (all mirror inputs exist) but not validated or tuned; pick the time step yourself (≈10 min at T127, ≈5 min at T255 as a start).

© climate-analytics-lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (scripts) in .claude/skills/derecho-jcm-runs of climate-analytics-lab/jax-gcm.

  • SKILL.md
  • reference/data_paths.md
  • scripts/mkjob.py
  • scripts/mkjob_test.py
  • scripts/settled_rate.py
  • scripts/watch_job.sh
  • scripts/watch_job_test.py

Open the folder on GitHubat commit 0940e89

Compare with similar skills

Derecho Jcm Runs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Derecho Jcm Runs compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Derecho Jcm Runs this skillclimate-analytics-lab/jax-gcm108—~2.5kAutomated safety check: PassApache-2.0
Megatron-LM on SLURMNVIDIA/Megatron-LM18k—~1.8kAutomated safety check: PassApache-2.0
Flash AttentionLuciole-Studio/Misaka-Agent1711 repos~2.7kAutomated safety check: PassMIT
Tensorflow Savedmodel Creatorjeremylongshore/tons-of-skills-marketplace2.8k—~593Automated safety check: PassMIT
Tensorflow Serving Setupjeremylongshore/tons-of-skills-marketplace2.8k—~578Automated safety check: PassMIT
Senior ML Engineerdavila7/claude-code-templates33k2 repos~1.4kAutomated safety check: PassMIT

Similar skills

  • Megatron-LM on SLURM

    NVIDIA/Megatron-LM

    Official

    Shows how to launch distributed Megatron-LM training on a SLURM cluster: sbatch skeleton, torch.distributed.run setup, CUDA_DEVICE_MAX_CONNECTIONS rules and failure diagnosis.

    18k GitHub stars~1.8k tokensUpdated today
    AI & LLM EngineeringAuto-check passed
  • Flash Attention

    Luciole-Studio/Misaka-Agent

    Speed up long-sequence transformer training and inference. An agent skill from Luciole-Studio/Misaka-Agent.

    171 GitHub starsUsed in 1 repo~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Tensorflow Savedmodel Creator

    jeremylongshore/tons-of-skills-marketplace

    Create tensorflow savedmodel creator operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~593 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Tensorflow Serving Setup

    jeremylongshore/tons-of-skills-marketplace

    Configure tensorflow serving setup operations. An agent skill from jeremylongshore/tons-of-skills-marketplace.

    2.8k GitHub stars~578 tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Senior ML Engineer

    davila7/claude-code-templates

    World-class ML engineering skill for productionizing ML models, MLOps, and building scalable ML systems.

    33k GitHub starsUsed in 2 repos~1.4k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

More from climate-analytics-lab/jax-gcm

  • Kubernetes Jcm Runs

    climate-analytics-lab/jax-gcm

    Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results.

    108 GitHub stars~3.4k tokensUpdated yesterday
    Auto-check passed
  • Devbox Jcm Runs

    climate-analytics-lab/jax-gcm

    Run jcm on the shared UCSD dev workstation (8x A100-80GB, no scheduler) — find a genuinely free GPU, avoid stomping on colleagues' jobs, environment and scratch paths, and the etiquette/traps…

    108 GitHub stars~1.7k tokensUpdated yesterday
    Auto-check passed
  • Jcm Benchmark

    climate-analytics-lab/jax-gcm

    Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion.

    108 GitHub stars~3.7k tokensUpdated yesterday
    Auto-check passed
  • Jcm Dev Workflow

    climate-analytics-lab/jax-gcm

    End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing…

    108 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Jcm Local CI

    climate-analytics-lab/jax-gcm

    Run the jax-gcm CI gates locally on Derecho when GitHub Actions minutes are exhausted or a pre-push check is wanted — lint, fast tests (90% coverage), slow tests (80% PR coverage) and a local Claude…

    108 GitHub stars~3k tokensUpdated yesterday
    Auto-check passed
  • Jcm Run

    climate-analytics-lab/jax-gcm

    Launch a jcm model run through the built-in Hydra configs — config groups, the validated stable T63L47 overrides, Hydra override traps, and watching for startup failures.

    108 GitHub stars~2.1k tokensUpdated yesterday
    Auto-check passed

Questions about Derecho Jcm Runs

What does Derecho Jcm Runs do?

Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues. Derecho Jcm Runs is an agent skill from climate-analytics-lab/jax-gcm. Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues.

When should I use Derecho Jcm Runs?

Derecho Jcm Runs fits situations like: running any jcm model integration; timestep/resolution sweep; performance benchmark; GPU job on Derecho — covers job-script generation.

How do I install Derecho Jcm Runs in Claude Code?

Run `npx skills add climate-analytics-lab/jax-gcm --skill derecho-jcm-runs -a claude-code`. Or copy the skill folder (.claude/skills/derecho-jcm-runs in climate-analytics-lab/jax-gcm) into .claude/skills/derecho-jcm-runs in your project. Claude Code loads it when a task matches its description.

How do I install Derecho Jcm Runs in Codex?

Run `npx skills add climate-analytics-lab/jax-gcm --skill derecho-jcm-runs -a codex`. Or copy the skill folder (.claude/skills/derecho-jcm-runs in climate-analytics-lab/jax-gcm) into .agents/skills/derecho-jcm-runs in your project. Codex loads it when a task matches its description.

Can I use Derecho Jcm Runs in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add climate-analytics-lab/jax-gcm --skill derecho-jcm-runs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/derecho-jcm-runs, .gemini/skills/derecho-jcm-runs, .github/skills/derecho-jcm-runs and .opencode/skills/derecho-jcm-runs in your project.

What does Derecho Jcm Runs need to run?

Going by SKILL.md and its folder, Derecho Jcm Runs needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (python). Our summary lists: Python 3; A Bash shell.

Does Derecho Jcm Runs access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Derecho Jcm Runs safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.

What licence does Derecho Jcm Runs use?

Derecho Jcm Runs is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Derecho Jcm Runs use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Derecho Jcm Runs?

Skills that share tags, products or a category with Derecho Jcm Runs: Megatron-LM on SLURM (NVIDIA/Megatron-LM, 18k stars), Flash Attention (Luciole-Studio/Misaka-Agent, 171 stars), Tensorflow Savedmodel Creator (jeremylongshore/tons-of-skills-marketplace, 2.8k stars) and Tensorflow Serving Setup (jeremylongshore/tons-of-skills-marketplace, 2.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Derecho Jcm Runs?

climate-analytics-lab (a GitHub organization) maintains it in climate-analytics-lab/jax-gcm, which has 108 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 10, 2026.

Source: climate-analytics-lab/jax-gcm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.