Dstack
dstackai/dstack
dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.
Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results.
$ npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a claude-codeProject install by default; add -g for ~/.claude/skills/.
$ gh skill install climate-analytics-lab/jax-gcm kubernetes-jcm-runs --agent claude-codeProject scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/kubernetes-jcm-runs .claude/skills/kubernetes-jcm-runs && rm -rf skills-srcUse ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.
Claude Code skills documentation · loads skills from .claude/skills/
Install the "kubernetes-jcm-runs" agent skill from https://github.com/climate-analytics-lab/jax-gcm/tree/dev/.claude/skills/kubernetes-jcm-runs into .claude/skills/kubernetes-jcm-runs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kubernetes-jcm-runs", then confirm the skill loads.Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$skill-installer install https://github.com/climate-analytics-lab/jax-gcm/tree/dev/.claude/skills/kubernetes-jcm-runsType this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
$ npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a codexProject install goes to .agents/skills/; add -g for ~/.codex/skills/.
$ gh skill install climate-analytics-lab/jax-gcm kubernetes-jcm-runs --agent codexProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .agents/skills && cp -r skills-src/.claude/skills/kubernetes-jcm-runs .agents/skills/kubernetes-jcm-runs && rm -rf skills-srcUse ~/.agents/skills/ instead of .agents/skills for a personal install.
Codex skills documentation · loads skills from .agents/skills/
Install the "kubernetes-jcm-runs" agent skill from https://github.com/climate-analytics-lab/jax-gcm/tree/dev/.claude/skills/kubernetes-jcm-runs into .agents/skills/kubernetes-jcm-runs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kubernetes-jcm-runs", then confirm the skill loads.Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a cursorProject install goes to .agents/skills/; add -g for ~/.cursor/skills/.
$ gh skill install climate-analytics-lab/jax-gcm kubernetes-jcm-runs --agent cursorProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .cursor/skills && cp -r skills-src/.claude/skills/kubernetes-jcm-runs .cursor/skills/kubernetes-jcm-runs && rm -rf skills-srcUse ~/.cursor/skills/ instead of .cursor/skills for a personal install.
Cursor skills documentation · loads skills from .cursor/skills/, .agents/skills/, .claude/skills/, .codex/skills/
Install the "kubernetes-jcm-runs" agent skill from https://github.com/climate-analytics-lab/jax-gcm/tree/dev/.claude/skills/kubernetes-jcm-runs into .cursor/skills/kubernetes-jcm-runs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kubernetes-jcm-runs", then confirm the skill loads.Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gemini skills install https://github.com/climate-analytics-lab/jax-gcm.git --path .claude/skills/kubernetes-jcm-runs--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
$ npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a gemini-cliProject install goes to .agents/skills/; add -g for ~/.gemini/skills/.
$ gh skill install climate-analytics-lab/jax-gcm kubernetes-jcm-runs --agent gemini-cliProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .gemini/skills && cp -r skills-src/.claude/skills/kubernetes-jcm-runs .gemini/skills/kubernetes-jcm-runs && rm -rf skills-srcUse ~/.gemini/skills/ instead of .gemini/skills for a personal install, then run /skills reload.
Gemini CLI skills documentation · loads skills from .gemini/skills/, .agents/skills/
Install the "kubernetes-jcm-runs" agent skill from https://github.com/climate-analytics-lab/jax-gcm/tree/dev/.claude/skills/kubernetes-jcm-runs into .gemini/skills/kubernetes-jcm-runs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kubernetes-jcm-runs", then confirm the skill loads.Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ gh skill install climate-analytics-lab/jax-gcm kubernetes-jcm-runsInstalls for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
$ npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a github-copilotProject install goes to .agents/skills/; add -g for ~/.copilot/skills/.
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .github/skills && cp -r skills-src/.claude/skills/kubernetes-jcm-runs .github/skills/kubernetes-jcm-runs && rm -rf skills-srcUse ~/.copilot/skills/ instead of .github/skills for a personal install. Commit .github/skills so cloud agent and code review can use it.
GitHub Copilot skills documentation · loads skills from .github/skills/, .claude/skills/, .agents/skills/
Install the "kubernetes-jcm-runs" agent skill from https://github.com/climate-analytics-lab/jax-gcm/tree/dev/.claude/skills/kubernetes-jcm-runs into .github/skills/kubernetes-jcm-runs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kubernetes-jcm-runs", then confirm the skill loads.GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
$ npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a opencodeOpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
$ gh skill install climate-analytics-lab/jax-gcm kubernetes-jcm-runs --agent opencodeProject scope by default (.agents/skills/); add --scope user for a personal install.
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .opencode/skills && cp -r skills-src/.claude/skills/kubernetes-jcm-runs .opencode/skills/kubernetes-jcm-runs && rm -rf skills-srcUse ~/.config/opencode/skills/ instead of .opencode/skills for a personal install.
OpenCode skills documentation · loads skills from .opencode/skills/, .claude/skills/, .agents/skills/
Install the "kubernetes-jcm-runs" agent skill from https://github.com/climate-analytics-lab/jax-gcm/tree/dev/.claude/skills/kubernetes-jcm-runs into .opencode/skills/kubernetes-jcm-runs/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "kubernetes-jcm-runs", then confirm the skill loads.OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
kubernetes-jcm-runsRun jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results.
Kubernetes Jcm Runs is an agent skill from climate-analytics-lab/jax-gcm. Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results. Currently configured for Nautilus/NRP; other clusters are a site profile in sites.py. Use for runs that need dedicated GPUs, especially when the shared dev workstation is saturated.
Its SKILL.md is about 3.4k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including scripts (for example `scripts/fetch_reports.py`, `scripts/fetch_run.py` and `scripts/mkjob.py`).
It sits in DevOps & Cloud, covering Container orchestration and GPU and accelerator computing. It works with Kubernetes. The repository describes itself as: GCM Physics written in JAX. The licence is Apache-2.0.
Read from SKILL.md and the folder at commit 0940e89. It shows what the files ask for, not the result of running them.
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Ships 6 files in scripts/ (Python and Shell), which the agent can run.
Shell commands in SKILL.md call:
kubectlpythonFrom the folder's file list and the shell code blocks in SKILL.md.
Hosts in commands or code, which the agent is likely to contact:
authentik.nrp-nautilus.ioAlso links to:
github.comFrom URLs in SKILL.md, links to its own repository left out.
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Kubernetes Jcm Runs loads about 3.4k tokens when it runs. Until then it costs about 87 tokens; SKILL.md has 1,727 words of instructions outside code blocks.
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); the scripts in this folder are not scanned.
The full file from climate-analytics-lab/jax-gcm at commit 0940e89, republished under its Apache-2.0 licence (© climate-analytics-lab). 1,727 words, ~3,448 tokens.
.claude/skills/kubernetes-jcm-runs/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.The Kubernetes site layer. Read jcm-run for the model/Hydra layer and
jcm-benchmark for throughput methodology and reference numbers; this file
covers only what running on a cluster adds.
Why use a cluster. Kubernetes gives a pod exclusive use of the GPUs it
requests. That is the guarantee benchmarking needs and the one a shared
workstation cannot offer — there, a neighbour can land on your card mid-run
and quietly corrupt the timing (devbox-jcm-runs). Quota permitting, a
multi-config sweep also runs in parallel, finishing in the time of its
slowest member rather than the sum.
Portability. The Job shape — clone pinned SHAs, install what the image
lacks, run jcm, write to a volume, refuse or survive failure — is generic.
Everything cluster-specific lives in scripts/sites.py as a profile:
namespace, GPU resource name and selector, storage class, PVC names, image,
and the extra pip packages. A second cluster is a new dict there plus
--site <name>, not a fork of the generators. Only Nautilus is defined;
sites.py documents what a GKE/EKS profile would have to supply (most
notably nvidia.com/gpu rather than a vendor quota bucket).
S=.claude/skills/kubernetes-jcm-runs/scripts
python $S/mkjob.py --preset ma-t63-l47 --f32 # inspect first
python $S/mkjob.py --preset ma-t63-l47 --f32 | kubectl apply -f -
python $S/mkjob.py --sweep --f32 | kubectl apply -f - # parallel sweepAlways look at the manifest before applying it. kubectl apply --dry-run=server -f - validates against the API without creating anything.
| flag | why |
|---|---|
--suffix TAG | Jobs are immutable, so a rerun collides with the previous Job (TTL 24 h). The suffix goes into the Job name and the report label, so the two results sit side by side instead of the second overwriting the first. |
--save-interval N | Days between writes. Defaults to --chunk-days. |
--env K=V | Container env var, repeatable — for knobs that are env-only rather than Hydra keys. |
--gpu-product P | Pin an exact GPU product; by default any card the site selector admits. |
--site NAME | Site profile from sites.py. |
--sweep skips presets whose inputs no pod can reach, and says which and
why. That check is derived from the preset's own file overrides, so a preset
becomes runnable automatically once its data is reachable rather than
needing an exclusion list edited.
pip install -e . and no extras, so anything newer or optional is added at pod
start (see extra_pip in the site profile, plus per-preset additions).
A hard GPU check follows every install, because a pip-induced CPU
fallback would still "work" and report timings ~100× slow.backoffLimit: 0 for benchmarks. A failed benchmark must be
inspected, not silently retried onto a different node with different
timings.emptyDir that dies with the pod; only reports and
logs reach the PVC. A benchmark needs the timings, not the fields./dev/shm as a memory-backed emptyDir. The 64 MB Kubernetes default
is far too small for JAX/XLA.--gpu 0 — the pod sees exactly one GPU, so the harness's free-GPU
gate is a no-op. Kubernetes has already guaranteed exclusivity, which is
the whole reason to run here.hf:// paths lazily during model construction, which is after
the telemetry sampler starts, so a multi-GB bundle would otherwise
download inside the timed region.mkrun.pyBenchmarks and production runs want opposite things, so they have separate
generators. Do not use mkjob.py for a run whose output you intend to keep.
python $S/mkrun.py --name pd-year | kubectl apply -f - # 12 calendar months
python $S/mkrun.py --name aci-2yr --months 24 --physics echam-jam-aci \
--pin jcm=abc1234 | kubectl apply -f -A run is 12 calendar months from --start-time (default 2000-01-01) by
default, under the present-day climatological AMIP forcing (the mirror's
forcing_pd/ozone_pd bundles, auto PD emissions/oxidants), written as
calendar-month means (<name>_monthly_YYYY-MM.nc) from daily means in 5-day
chunks. --days N gives a fixed length instead; --save-chunks also keeps
the per-chunk daily files (hundreds of GB for a JAM year). The completion
gate reads the per-chunk Chunk K | Day N health lines, so it works without
chunk files.
mkjob.py (benchmark) | mkrun.py (production) | |
|---|---|---|
| model output | pod-local emptyDir, discarded | runs PVC, kept |
| health gate | --allow-unhealthy optional | always on — stop on NaN |
backoffLimit | 0 — inspect a failure | 20 — survive eviction |
| on retry | would re-time on another node | resumes from checkpoint |
| TTL | 24 h | none |
Eviction tolerance is the whole design. A pod is not guaranteed a node
for days, so a long run must assume interruption. jcm resumes automatically
when its checkpoint exists, so a restarted pod picks up at the last
completed chunk instead of restarting the year — that is what makes
multi-day runs viable here at all. --retries is effectively the eviction
budget.
Pin the code for anything long: --pin jcm=<sha>. A run resumed days later
must come back on the same code; a branch name would silently restart it
on whatever has since been merged, mid-year.
Output lands under /runs/<name>/ alongside a PROVENANCE file recording
the SHA of every repo, appended on each attempt so a resumed run shows its
full history.
A run that stops early can still exit 0. run_chunked returns normally
when the health gate trips, so Kubernetes would mark a truncated year
Complete. mkrun.py therefore checks the health verdict and the day
count reached, and fails the Job if either falls short.
Do not hand-build these from mkrun.py's output. The release matrix
(tools/release_validation/matrix.yaml) and one-member arms of it (the
#682 retune) go through the release launcher, which builds the same Job
through mkrun.job_manifest() — the importable engine behind mkrun.py —
from the matrix's own override list:
python tools/release_validation/launch.py --site nautilus --submit # 7 members
python tools/release_validation/launch.py --site nautilus --members echam-jam-t63-l47 \
--days 60 --init <state> --suffix entrpen20 \
--extra +physics.convection.entrpen=2e-4 --submit # one arm
python tools/release_validation/launch.py --site nautilus --fetch --members <m> --tag <t>It pins jcm to a pushed SHA and the image to its digest, installs that
commit's own requirements in the pod (locked on the first attempt), records
each launch so --resume re-emits it unchanged, and has the pod refuse a run
directory that belongs to a different launch, or a fresh Job onto a run
another Job started (the PBS path's check_fresh, done where the volume is
visible). Workflow,
the arm recipe and the fetch/score/ingest commands:
tools/release_validation/README.md ("Workflow on Nautilus").
A new door that keeps a run's output should call mkrun.job_manifest() too,
not copy its YAML or rewrite its jcm.main line.
kubectl get pods -l job-name=<job>
kubectl logs -f -l job-name=<job>
kubectl get jobs # COMPLETIONS column
python $S/fetch_reports.py # summary table from job logs
python $S/fetch_reports.py --copy ./ # also copy them locally
python $S/fetch_reports.py --from-pvc # once a job's TTL has expiredfetch_reports.py reads job logs by default — no pod, no exec, no volume
mount. --from-pvc falls back to a throwaway pod that mounts the reports
volume, which is what you need after the 24 h TTL removes the Job.
A production run's directory comes off the runs volume with
fetch_run.py <run> <dest>: a read-only CPU reader pod and a tar stream
through kubectl exec, incremental and size-checked, checkpoints skipped
unless --with-checkpoints. The reader pod is always deleted.
A pod that vanished is not a pod that succeeded. Check the Job's
COMPLETIONS, and read the report — the harness refuses to quote a rate for
a truncated or NaN'd run, so a report with no throughput line means the run
failed even if Kubernetes says Completed.
kubectl cp is slow and has no resume, so do not pull multi-GB netCDF with
it; fetch_run.py (above) resumes per file, and for bulk well beyond a run's
chunk files object storage pushed from a pod is still the faster route.
The National Research Platform. Everything below is specific to it.
export PATH=/data/dwatsonparris/micromamba/bin:$PATH # kubectl lives here
kubectl config current-context # -> nautilus
kubectl get pods # namespace: climate-analyticsAuthentication is OIDC via an exec plugin, which needs
kubectl-oidc_login (int128/kubelogin)
on PATH. Note conda-forge's kubelogin package is Azure's tool and is
not the same thing.
The token caches under ~/.kube/cache/oidc-login/. When it expires, re-run
the device-code flow — interactive, cannot be automated:
kubectl oidc-login get-token --grant-type=device-code --skip-open-browser \
--oidc-issuer-url=https://authentik.nrp-nautilus.io/application/o/k8s/ \
--oidc-client-id=<client-id> --oidc-extra-scope=profile,offline_access--skip-open-browser matters: without it kubelogin tries to launch a
browser on a headless box and hangs.
nvidia.com/a100 is a quota bucket, not a card type. It spans four
products:
| product | nodes | memory |
|---|---|---|
| NVIDIA-A100-SXM4-80GB | 22 | 80 GB |
| NVIDIA-A100-80GB-PCIe | 9 | 80 GB |
| NVIDIA-A100-PCIE-40GB | 8 | 40 GB |
| NVIDIA-A100-80GB-PCIe-MIG-1g.10gb | 1 | 9.7 GB |
A bare nvidia.com/a100: 1 request can land on any of them. The 40 GB card
OOMs the larger configs and the MIG slice OOMs almost everything — and a
benchmark that lands on a different product each run produces numbers that
cannot be compared at all. mkjob.py therefore selects on the memory
label, keeping both 80 GB variants and excluding the other two:
nodeSelector:
nvidia.com/gpu.memory: "81920"PCIe vs SXM4 was measured, not assumed: 0.3 % apart on an identical
config, despite SXM4 being a 400 W part against PCIe's 300 W. jcm is memory-
and launch-bound rather than power-bound at these sizes, so the extra
headroom buys nothing. Do not spend scheduling flexibility pinning a product
without a reason; --gpu-product is there if a comparison genuinely needs
it, and every job records what it ran on (gpu_product.txt) so a surprising
number can be checked rather than re-litigated.
ghcr.io/climate-analytics-lab/jcm:latest — public, GPU-ready
(nvidia/cuda:12.6.3-base + jax[cuda12]). Tags include releases
(v2.0.1) and manual-<sha> builds.jcm-bench PVC — 50 Gi RWX rook-cephfs, so parallel pods share it;
holds reports and logs.jcm-runs PVC — 500 Gi RWX, production output.kubectl describe resourcequota a100-limit shows
current use; a ninth pod sits Pending rather than failing. Pod cap is
200, well above the GPU limit.kubectl delete job <name> first, or use --suffix.ttlSecondsAfterFinished: 86400 removes benchmark Jobs after a day.
Collect reports before then, or read the PVC with --from-pvc.priorityClassName. The namespace bans every named class
at 0 pods — including default — via the high-priority-ban and
low-priority-ban quotas. A pod that names one is refused by quota; an
unnamed pod runs at priority 0 and is fine. Both generators leave it unset
deliberately.kubectl get node is forbidden, so a pod's
GPU product must come from the pod itself, not the node label.jcm-benchmark — methodology and reference numbers; everything there applies herejcm-run — config groups and Hydra trapsdevbox-jcm-runs — the shared workstation, and why it is worse for timingderecho-jcm-runs — the PBS alternative (40 GB cards, so memory limits differ)© climate-analytics-lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file
SKILL.md and 6 other files (scripts) in .claude/skills/kubernetes-jcm-runs of climate-analytics-lab/jax-gcm.
Open the folder on GitHubat commit 0940e89
Kubernetes Jcm Runs next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
| Skill | Stars | Used in | Tokens | Auto-check | Licence | Repo updated |
|---|---|---|---|---|---|---|
| Kubernetes Jcm Runs this skillclimate-analytics-lab/jax-gcm | 108 | — | ~3.4k | Automated safety check: Pass | Apache-2.0 | |
| Dstackdstackai/dstack | 2.3k | — | ~6.2k | Automated safety check: Warn | MPL-2.0 | |
| GPU Kubernetes Operationssickn33/agentic-awesome-skills | 47k | 2 repos | ~3.2k | Automated safety check: Pass | MIT | |
| Paidf Orchestration SetupNVIDIA/skills | 3.6k | — | ~3.8k | Automated safety check: Warn | Apache-2.0 | |
| QzcliAI4Scientist/nano-scientist | 128 | 2 repos | ~1.9k | Automated safety check: Notes | None | |
| Kubeshark Installerkubeshark/kubeshark | 12k | — | ~3.6k | Automated safety check: Notes | Apache-2.0 |
dstackai/dstack
dstack is an open-source control plane for GPU provisioning and orchestration across GPU clouds, Kubernetes, and on-prem clusters.
sickn33/agentic-awesome-skills
Operate GPU-backed Kubernetes clusters for AI inference and training with scheduling, autoscaling, node health, MIG partitioning, and cost controls.
NVIDIA/skills
Audit, prepare, and deploy PAIDF Orchestration on a Kubernetes GPU cluster - single-GPU H100/L40S hosts, managed Kubernetes, kubeadm, and similar.
AI4Scientist/nano-scientist
Manage GPU compute jobs on the Qizhi (启智) platform using qzcli — a kubectl-style CLI tool.
kubeshark/kubeshark
Installs and configures Kubeshark on a Kubernetes cluster, choosing between the quick CLI path and a Helm install with custom values.
kubesphere/kubesphere
Creates and queries KubeSphere users, workspaces and projects and assigns built-in roles, defaulting to least privilege and never deleting anything.
climate-analytics-lab/jax-gcm
Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues.
climate-analytics-lab/jax-gcm
Run jcm on the shared UCSD dev workstation (8x A100-80GB, no scheduler) — find a genuinely free GPU, avoid stomping on colleagues' jobs, environment and scratch paths, and the etiquette/traps…
climate-analytics-lab/jax-gcm
Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion.
climate-analytics-lab/jax-gcm
End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing…
climate-analytics-lab/jax-gcm
Run the jax-gcm CI gates locally on Derecho when GitHub Actions minutes are exhausted or a pre-push check is wanted — lint, fast tests (90% coverage), slow tests (80% PR coverage) and a local Claude…
climate-analytics-lab/jax-gcm
Launch a jcm model run through the built-in Hydra configs — config groups, the validated stable T63L47 overrides, Hydra override traps, and watching for startup failures.
Works with
Categories
Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results. Kubernetes Jcm Runs is an agent skill from climate-analytics-lab/jax-gcm. Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results.
Kubernetes Jcm Runs fits situations like: runs that need dedicated GPUs; especially when the shared dev workstation is saturated.
Run `npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a claude-code`. Or copy the skill folder (.claude/skills/kubernetes-jcm-runs in climate-analytics-lab/jax-gcm) into .claude/skills/kubernetes-jcm-runs in your project. Claude Code loads it when a task matches its description.
Run `npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a codex`. Or copy the skill folder (.claude/skills/kubernetes-jcm-runs in climate-analytics-lab/jax-gcm) into .agents/skills/kubernetes-jcm-runs in your project. Codex loads it when a task matches its description.
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add climate-analytics-lab/jax-gcm --skill kubernetes-jcm-runs -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/kubernetes-jcm-runs, .gemini/skills/kubernetes-jcm-runs, .github/skills/kubernetes-jcm-runs and .opencode/skills/kubernetes-jcm-runs in your project.
Going by SKILL.md and its folder, Kubernetes Jcm Runs needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (kubectl and python). Our summary lists: Python 3; A Bash shell.
SKILL.md names 2 domains. In commands or code: authentik.nrp-nautilus.io; the agent is likely to contact it when it follows the instructions. As links in the text: github.com. This is read from the text; nothing was executed.
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. The check reads SKILL.md only: the scripts in the folder are not scanned, so read them before running anything.
Kubernetes Jcm Runs is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.
About 3.4k tokens (SKILL.md is roughly 14k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
Skills that share tags, products or a category with Kubernetes Jcm Runs: Dstack (dstackai/dstack, 2.3k stars), GPU Kubernetes Operations (sickn33/agentic-awesome-skills, 47k stars), Paidf Orchestration Setup (NVIDIA/skills, 3.6k stars) and Qzcli (AI4Scientist/nano-scientist, 128 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
climate-analytics-lab (a GitHub organization) maintains it in climate-analytics-lab/jax-gcm, which has 108 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 10, 2026.
Source: climate-analytics-lab/jax-gcm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.