Agent skill

Jcm Dev Workflow

by climate-analytics-lab in climate-analytics-lab/jax-gcm

End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing…

Apache-2.0Auto-check passedTesting & QA

Install Jcm Dev Workflow

skills CLI
$ npx skills add climate-analytics-lab/jax-gcm --skill jcm-dev-workflow -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install climate-analytics-lab/jax-gcm jcm-dev-workflow --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/jcm-dev-workflow .claude/skills/jcm-dev-workflow && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
jcm-dev-workflow
GitHub stars
108
Token cost
~3k tokens
SKILL.md length
1,655 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing…

  • Works in 8 steps: Work locally, committing atomically → The local gate — run this before every… → Adversarial self-review — before the… → …
  • Any code change destined for a PR
  • SKILL.md covers 1. Work locally, committing…, 2. The local gate — run this…, 3. Adversarial self-review —… and 4. Push and open the PR, plus 5 more sections
  • Calls gh, ruff and pytest

What it does

Jcm Dev Workflow is an agent skill from climate-analytics-lab/jax-gcm. End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing back for human review only once everything is green. Use for any code change destined for a PR.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Testing & QA, covering Linting and formatting. It works with pytest. The repository describes itself as: GCM Physics written in JAX. The licence is Apache-2.0.

When your agent uses it

  • Any code change destined for a PR
  • Tasks that involve Linting and formatting

Example prompts

  • “/jcm-dev-workflow”

Workflow steps

8 steps, taken from the step headings in SKILL.md.

  1. Work locally, committing atomically
  2. The local gate — run this before every push
  3. Adversarial self-review — before the first push, and before any re-push that changes code
  4. Push and open the PR
  5. Monitor CI and Codex — both, from one watcher
  6. Fix CI, respond to Codex
  7. Re-request review when the response is substantial
  8. Hand back to the human only when green

What it can do on your machine

Read from SKILL.md and the folder at commit 35dc199. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • gh
    • ruff
    • pytest
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use gh and git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Jcm Dev Workflow loads about 3k tokens when it runs. Until then it costs about 78 tokens; SKILL.md has 1,655 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~78
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from climate-analytics-lab/jax-gcm at commit 35dc199, republished under its Apache-2.0 licence (© climate-analytics-lab). 1,655 words, ~2,980 tokens.

Download SKILL.mdSave it as .claude/skills/jcm-dev-workflow/SKILL.md (or your agent's skills folder).
name
jcm-dev-workflow
description
End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI *and* the automatic Codex review, addressing feedback, and handing back for human review only once everything is green. Use for any code change destined for a PR.

Development workflow

The loop, in order. Each step exists because skipping it costs a CI cycle, a review round, or a wrong result.

branch → atomic commits → test + lint locally → adversarial self-review → push
   → PR (linked to issue) → monitor CI *and* Codex → fix/respond
   → re-review if substantial → ping human

1. Work locally, committing atomically

Branch off dev (not main — clean releases are merged to main and tagged). Never commit directly to dev or main.

One logical change per commit, each self-contained and passing tests on its own. A commit message says why, not what the diff already shows: the reference formulation being matched, the failure that motivated a guard, the alternative rejected and the reason. When a change encodes a scientific decision, that reasoning belongs in the commit and in a comment — see the "Think Before Coding" and documentation rules in CLAUDE.md.

Stacking several subtasks on one PR is fine and often easier to review than a chain of dependent PRs — say so in the PR body so the reviewer knows the commits are separable.

2. The local gate — run this before every push

bash
ruff check .                          # MUST be clean; it is the only linter
JAX_PLATFORMS=cpu pytest -n 12        # full suite, ~2 min
JAX_PLATFORMS=cpu pytest -n 12 -m "not slow"   # fast subset while iterating

JAX_PLATFORMS=cpu is required on this GPU host. Without it every xdist worker grabs the same GPU and you get CUDA_ERROR_OUT_OF_MEMORY / dnn_support != nullptr RET_CHECK failures from XLA. The unit tests are small column-mode integrations that run faster on CPU than they would round-tripping through the device.

  • -n 12 is the local default; -n auto picks from visible CPUs; -n 0 (or omitting -n) forces one process when you need ordered output or are chasing a flake.
  • Coverage: JAX_PLATFORMS=cpu pytest -n 12 --cov=jcm --cov-fail-under=90, then coverage report --fail-under=90. The second command is the one that reliably fails: fail_under is judged at the reported precision, which is why both rcfiles set [report] precision = 2 (#786).
  • Tests are *_test.py, co-located with the module, unittest.TestCase run under pytest. Root conftest.py clears jcm imports between tests to stop state leaking.
  • Mark tests over ~1 min @pytest.mark.slow.

CI thresholds: ruff gates the run, then the fast tests at 90% coverage and the slow tests at 80% (pull requests only) run in parallel behind it — the slow suite as two path shards (slow-tests-radiation, slow-tests-rest, defined in tools/ci/slow_shards.py) whose coverage the slow-coverage job combines before enforcing the floor. A separate extras-tests job installs every optional extra and runs the tests they gate (JCM_REQUIRE_EXTRAS=1 pytest -m requires_extra, no coverage floor), on every PR and push to main/dev; gate a test on an extra only with @pytest.mark.requires_extra(...), since every job fails any other gate. On a PR, if the fast suite goes red the run is cancelled, taking the slow suite and extras-tests with it, and fast-tests itself reports as cancelled rather than failed (the failing step is still red inside it). A cancelled slow result therefore never means passing — and never means the fast suite failed either, since cancel-in-progress cancels it the same way when your next push supersedes the run. Open fast-tests before concluding anything. push triggers the workflow on main/dev alone, so a feature branch gets no CI until its PR exists: run the full suite locally before opening one, or the PR is the first thing that has ever tested it.

Lint before every push, always. ruff check . takes seconds; a lint failure in CI burns a full cycle on something reported instantly locally. Treat it as part of the definition of done.

Beyond the suite: a change to physics or numerics needs a representative run, not just green tests. Verify conservation/physical correctness, and for anything performance-related use jcm-benchmark — never quote the cumulative sim days/hr line.

3. Adversarial self-review — before the first push, and before any re-push that changes code

Codex reviews every push and its credits are finite. A finding Codex makes that a local reviewer would have made is a credit burnt and a review round lost (hours, with CI in the loop). So before git push, on the branch:

/code-review high        # multi-angle finders + one verifier per finding, diff vs upstream

Only from an Opus or Sonnet session. /code-review forks the invoking session on the same model with its full context, and fans out the same way; from a Fable session it spends Fable tokens at full context and has exhausted a session limit. On a Fable session run the same review as an explicit model: opus general-purpose agent that posts one gh pr review --comment (recipe in jcm-local-ci, "Local Claude review").

Treat the output exactly like a Codex review (step 6): fix every CONFIRMED finding, sweep for the same mistake elsewhere, and for anything you decide not to change record why in the commit or PR body — a finding refuted once locally is not re-argued with Codex later. Every fix the review produces goes back through step 2 — ruff check . and the tests — before the push: a review-suggested change is untested code until it does. If the fixes were more than trivial, run the review again too. Only then push. If a diff must go up before the review is clean (to share it, to reach GPU CI), open the PR as a draft and say so in the body; the rule still applies to the next push.

Scope: any change to executable code, wherever it lives — jcm/, tools/, utils/, validation/, docs/generate_docs.py, .claude/skills/*/scripts, workflows, root scripts such as run_benchmarks.sh; the list is illustrative, not a boundary. Docs-only and comment-only changes are exempt.

4. Push and open the PR

bash
git push -u origin <branch>
gh pr create --base dev --title "..." --body "..."

Link the issue so it closes on merge: Closes #NNN (or Refs #NNN for partial work) in the body. If the PR covers a subtask of a tracking issue, say which.

Write the body for the reviewer: what changed, why this approach over the alternatives, the decision anyone would second-guess, measured cost for anything performance-affecting, and what you verified. Implementation-specific detail and gotchas go here; general design goes in docs/source/design/ per CLAUDE.md.

5. Monitor CI and Codex — both, from one watcher

A Codex review arrives automatically on push; it is not triggered by you. Watch for both, because green CI with unaddressed Codex findings is not done.

bash
# CI checks
gh pr checks <PR> --watch

# Codex posts as a REVIEW with inline comments, not as an issue comment --
# `gh pr view --json comments` will NOT show it.
gh api repos/climate-analytics-lab/jax-gcm/pulls/<PR>/reviews \
  -q '.[] | "\(.user.login) | \(.state) | \(.submitted_at)"'
gh api repos/climate-analytics-lab/jax-gcm/pulls/<PR>/comments \
  -q '.[] | "### \(.path):\(.line)\n\(.body)\n"'

Arm a persistent Monitor that polls both and emits a line per new check result or new review comment, exiting when checks are complete and a review has landed. Make the filter cover failure signatures too — a watcher that matches only success is silent through a crash, which looks exactly like "still running".

Show full SKILL.md (634 more words)Show less

6. Fix CI, respond to Codex

Fix every CI failure. Do not re-push hoping it was a flake; if you believe it is one, say why.

For each Codex finding: verify it against the code before acting. They are usually right and often catch real defects, but not always — confirm the failure it describes is reachable. Verify against the code, not against your notes or memory: a finding that contradicts something you believe is the highest-value kind, because one of the two is stale and it may not be the finding. Then either fix it, or reply explaining concretely why it does not apply. Never silently ignore one.

Findings are labelled P1/P2/P3. P1s block. Reply on the PR summarising what you changed for each, so the human reviewer can see the loop closed without re-deriving it.

Fix the class, not the instance. Before calling a finding done, grep for the same mistake elsewhere — a reviewer reports the instance they happened to look at, not every occurrence. This is the single most common way a review round gets repeated: in one session the same review loop re-reported a NaN-check hole fixed in one file and left in another, and an sys.executable-vs-run-interpreter assumption fixed in one script and left in its sibling. Both were avoidable with one grep.

Enumerate the category; do not grep for the forms you were shown. This is the part that is easy to get wrong, and getting it wrong looks exactly like having done the sweep. Grepping for the specific strings in the review comment finds only what you already knew about.

The same session ran three rounds of one defect — diagnostics reading step-start state under operator splitting — because each sweep searched the reported spellings (state.tracers, then clouds.qc) instead of asking "which prognostic fields does any diagnostic read?". The third round finally enumerated every state.<field> access in every diagnostics-category term and produced a table with a verdict per site. That table is the deliverable; a grep hit-count is not.

So: name the invariant ("a diagnostic must report the state as saved"), list every place that invariant could be violated, and check each — including the ones that turn out fine, because "already correct" is a result worth recording. If your fix added a mechanism, check you consumed all of it: in that session the new accumulator already carried the wind and humidity tendencies that the next round's findings were about.

Check the invariant at each USE, not at each textual read. The same session's fourth round was a value that arrived as a parameter, three lines from a call site fixed two commits earlier — no grep over state.<field> could have found it, because the source text at that point said temperature. Enumerating reads finds direct accesses; it misses anything passed in, aliased, or defaulted. When the invariant is about provenance ("is this value the post-physics one?"), follow the values, not the spellings.

7. Re-request review when the response is substantial

If the fixes were more than trivial, ask for another pass with a comment containing exactly:

@codex review

Then monitor again (step 5). Repeat until a review comes back with nothing material.

8. Hand back to the human only when green

Ping the user for review once all of:

  • CI checks pass
  • Codex has no outstanding material findings
  • ruff check . is clean and the full suite passes locally
  • docs updated in the same PR if user-facing behaviour changed (CLAUDE.md: documentation lives with the change)

Then summarise: what landed, what was measured, what you decided and why, and anything you are still unsure about. Flag open questions explicitly rather than letting them pass as settled — an unflagged uncertainty is worse than a known gap.

  • jcm-run — driving the model (config groups, Hydra traps)
  • jcm-benchmark — measuring throughput without fooling yourself
  • devbox-jcm-runs / derecho-jcm-runs / kubernetes-jcm-runs — machine-specific execution

© climate-analytics-lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/jcm-dev-workflow of climate-analytics-lab/jax-gcm.

Open the folder on GitHubat commit 35dc199

Compare with similar skills

Jcm Dev Workflow next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Jcm Dev Workflow compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Jcm Dev Workflow this skillclimate-analytics-lab/jax-gcm108—~3kAutomated safety check: PassApache-2.0
Simple Modern Uvjlevy/simple-modern-uv301—~1.9kAutomated safety check: PassMIT
Running Unit TestsPSBrew/MkPFS169—~429Automated safety check: PassGPL-3.0
Gateguardana/guardana152—~776Automated safety check: PassApache-2.0
Python Helpershepherdjerred/monorepo112—~2.3kAutomated safety check: PassGPL-3.0
Kedro Babysitkedro-org/kedro11k—~4kAutomated safety check: PassCustom licence

Similar skills

  • Simple Modern Uv

    jlevy/simple-modern-uv

    Start, selectively modernize, fully migrate, or update Python projects using simple-modern-uv practices: uv, ruff, BasedPyright, pytest, GitHub Actions CI, and tag-driven PyPI publishing.

    301 GitHub stars~1.9k tokensUpdated 1 mo ago
    Testing & QAAuto-check passed
  • Running Unit Tests

    PSBrew/MkPFS

    Run formatting, linting, and tests; then iteratively fix failures until the branch is green.

    169 GitHub stars~429 tokensUpdated 4 days ago
    Testing & QAAuto-check passed
  • Gate

    guardana/guardana

    Run this project's verification — the full local CI mirror (ruff, mypy, import contract, pytest with PostgreSQL, coverage floors, dogfood, generated docs and site, the isolated example suites, the…

    152 GitHub stars~776 tokensUpdated yesterday
    Testing & QAAuto-check passed
  • Python Helper

    shepherdjerred/monorepo

    Current Python development guidance for versions, uv and pip, packaging, typing, asyncio, pytest, Ruff, security, and runtime boundaries.

    112 GitHub stars~2.3k tokensUpdated today
    Testing & QAAuto-check passed
  • Kedro Babysit

    kedro-org/kedro

    Run Kedro's local lint / format / type-check / tests on changed files (uses the project's pre-commit hooks, ruff, mypy, pytest, lint-imports, detect-secrets, Make targets — in the right venv), or…

    11k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Adk Setup

    google/adk-python

    Official

    Sets up a local ADK Python development environment in a git clone of the open-source adk-python repository: a uv virtual environment, all dependency extras, pre-commit hooks, and a first unit-test…

    22k GitHub stars~993 tokensUpdated today
    DevelopmentAuto-check: notes

More from climate-analytics-lab/jax-gcm

  • Derecho Jcm Runs

    climate-analytics-lab/jax-gcm

    Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues.

    108 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Kubernetes Jcm Runs

    climate-analytics-lab/jax-gcm

    Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results.

    108 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Devbox Jcm Runs

    climate-analytics-lab/jax-gcm

    Run jcm on the shared UCSD dev workstation (8x A100-80GB, no scheduler) — find a genuinely free GPU, avoid stomping on colleagues' jobs, environment and scratch paths, and the etiquette/traps…

    108 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Jcm Benchmark

    climate-analytics-lab/jax-gcm

    Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion.

    108 GitHub stars~3.7k tokensUpdated today
    Auto-check passed
  • Jcm Local CI

    climate-analytics-lab/jax-gcm

    Run the jax-gcm CI gates locally on Derecho when GitHub Actions minutes are exhausted or a pre-push check is wanted — lint, fast tests (90% coverage), slow tests (80% PR coverage) and a local Claude…

    108 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Jcm Run

    climate-analytics-lab/jax-gcm

    Launch a jcm model run through the built-in Hydra configs — config groups, the validated stable T63L47 overrides, Hydra override traps, and watching for startup failures.

    108 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Works with

Questions about Jcm Dev Workflow

What does Jcm Dev Workflow do?

End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing…. Jcm Dev Workflow is an agent skill from climate-analytics-lab/jax-gcm. End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing back for human review only once everything is green.

When should I use Jcm Dev Workflow?

Jcm Dev Workflow fits situations like: any code change destined for a PR; tasks that involve Linting and formatting.

How do I install Jcm Dev Workflow in Claude Code?

Run `npx skills add climate-analytics-lab/jax-gcm --skill jcm-dev-workflow -a claude-code`. Or copy the skill folder (.claude/skills/jcm-dev-workflow in climate-analytics-lab/jax-gcm) into .claude/skills/jcm-dev-workflow in your project. Claude Code loads it when a task matches its description.

How do I install Jcm Dev Workflow in Codex?

Run `npx skills add climate-analytics-lab/jax-gcm --skill jcm-dev-workflow -a codex`. Or copy the skill folder (.claude/skills/jcm-dev-workflow in climate-analytics-lab/jax-gcm) into .agents/skills/jcm-dev-workflow in your project. Codex loads it when a task matches its description.

Can I use Jcm Dev Workflow in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add climate-analytics-lab/jax-gcm --skill jcm-dev-workflow -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/jcm-dev-workflow, .gemini/skills/jcm-dev-workflow, .github/skills/jcm-dev-workflow and .opencode/skills/jcm-dev-workflow in your project.

What does Jcm Dev Workflow need to run?

Going by SKILL.md and its folder, Jcm Dev Workflow needs the command-line tools its instructions call (gh, ruff, pytest and git).

Does Jcm Dev Workflow access the network?

SKILL.md contains no URLs. Its commands use gh and git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Jcm Dev Workflow safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Jcm Dev Workflow use?

Jcm Dev Workflow is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Jcm Dev Workflow use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Jcm Dev Workflow?

Skills that share tags, products or a category with Jcm Dev Workflow: Simple Modern Uv (jlevy/simple-modern-uv, 301 stars), Running Unit Tests (PSBrew/MkPFS, 169 stars), Gate (guardana/guardana, 152 stars) and Python Helper (shepherdjerred/monorepo, 112 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Jcm Dev Workflow?

climate-analytics-lab (a GitHub organization) maintains it in climate-analytics-lab/jax-gcm, which has 108 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.

Source: climate-analytics-lab/jax-gcm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.