Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion.

Apache-2.0Auto-check passedAI & LLM Engineering

Install Jcm Benchmark

skills CLI
$ npx skills add climate-analytics-lab/jax-gcm --skill jcm-benchmark -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install climate-analytics-lab/jax-gcm jcm-benchmark --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/climate-analytics-lab/jax-gcm.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.claude/skills/jcm-benchmark .claude/skills/jcm-benchmark && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
jcm-benchmark
GitHub stars
108
Token cost
~3.7k tokens
SKILL.md length
2,014 words
Files
1
Skills in repo
7
Repo updated
First seen
Licence
Apache-2.0

At a glance

Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion.

  • Comparing hardware
  • SKILL.md covers Run it, "Which term is it?" —…, Methodology — the part that… and Configuration, plus 3 more sections
  • Calls python and git
  • Library versions (jax-rrtmgp

What it does

Jcm Benchmark is an agent skill from climate-analytics-lab/jax-gcm. Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion. Use when comparing hardware, library versions (jax-rrtmgp, dinosaur), physics presets, or checking for a performance regression.

Its SKILL.md is about 3.7k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in AI & LLM Engineering, covering Deep learning. The repository describes itself as: GCM Physics written in JAX. The licence is Apache-2.0.

When your agent uses it

  • Comparing hardware
  • Library versions (jax-rrtmgp
  • Physics presets
  • Checking for a performance regression

Example prompts

  • “/jcm-benchmark”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 20ca89d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • python
    • git

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use git, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Jcm Benchmark loads about 3.7k tokens when it runs. Until then it costs about 81 tokens; SKILL.md has 2,014 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~81
When it runs · the whole SKILL.md, loaded when a task matches
~3.7k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from climate-analytics-lab/jax-gcm at commit 20ca89d, republished under its Apache-2.0 licence (© climate-analytics-lab). 2,014 words, ~3,704 tokens.

Download SKILL.mdSave it as .claude/skills/jcm-benchmark/SKILL.md (or your agent's skills folder).
name
jcm-benchmark
description
Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion. Use when comparing hardware, library versions (jax-rrtmgp, dinosaur), physics presets, or checking for a performance regression.

Benchmarking jcm

Builds on jcm-run (config groups, Hydra traps) and the machine skill for your site — devbox-jcm-runs or derecho-jcm-runs. This skill is about getting a throughput number that is actually true, which is harder than it looks.

Run it

bash
PY=/home/dwatsonparris/micromamba/envs/jcm/bin/python

# Short: 1 month (30 d), ~6 chunks. The default for any A/B.
$PY tools/benchmark.py --preset t63-echam-rrtmgp --months 1 --gpu 1 \
    --label baseline

# Long: 12 months. Use for provisioning and drift, not for A/B.
$PY tools/benchmark.py --preset t63-echam-jam --months 12 --gpu 3 \
    --chunk-days 30 --label jam-year

# A/B a library version without touching the shared editable install
$PY tools/benchmark.py --preset t63-echam-rrtmgp --months 1 --gpu 1 \
    --label rrtmgp-perf \
    --pythonpath /data/dwatsonparris/jax-rrtmgp-worktrees/perf

Presets are in tools/benchmark.py:PRESETS. Each carries the validated stable override set for its grid, not just physics=/grid= — see "Configuration" below. Results land in /scr/dwatsonparris/benchmarks/<label>/ as report.md, result.json, run.log and gpu.csv.

Runs take tens of minutes; launch with run_in_background and watch with a Monitor on run.log, or the tool call will time out.

"Which term is it?" — tools/profile_terms.py

benchmark.py answers how fast; it cannot say where the time went, because a step is one fused XLA module. For that, same presets, different tool:

bash
$PY tools/profile_terms.py --preset ma-t63-l47 --gpu 3

It prints ms/step and ms/call for the dynamical core, each bridge direction and every physics term, by joining a profiler trace to the jax.named_scope labels that jcm.profiling puts in the HLO. Three rules when quoting it:

  • It disables CUDA graph capture so that kernels stay individually attributable, so its total step time is high by design (264 vs 236 ms at T63L47). Never quote it as throughput — that is benchmark.py's number.
  • Read the "mixed kernels" percentage first. That is device time in fusions spanning two components, charged whole to one of them; when it is large the per-term split is soft and only the coarse dynamics/physics/bridge division is safe to quote.
  • Do not lengthen the window to "get a better average". The profiler's event buffer holds ~1e6 events and a T63L47 JAM step emits ~19,000 kernels, so ~20 steps is the ceiling there. A one-day window recorded 20 of its 120 steps and reported a 4× undercount. The tool now fails instead, but the instinct is the trap. It is also unnecessary: kernel shapes are static and there are no data-dependent branches, so a step costs the same regardless of state — the window only needs to span whole radiation sub-cycles, which the default does.

Methodology — the part that matters

Never quote jcm's N sim days/hr log line. It is cumulative including compile time, so it understates throughput by 2-5x early and drifts upward for the whole run. A 5.3x regression was once filed against jax-rrtmgp on the strength of chunk 1 of a run that settled 22x faster; it had to be retracted. tools/benchmark.py ignores that line and parses Wall: Xs this chunk. The reduction lives in tools/chunk_timing.py, shared with the Derecho skill's settled_rate.py so a run cannot get two different answers depending on which tool read it.

Discard chunk 1 — it contains compilation. For T63L47 ECHAM+JAM, compile alone is several minutes.

Require convergence. Chunk times settle over 3-4 chunks (XLA autotuning, cache warming, allocator). A rate is quoted only once the last two chunks agree within 3% — judged on the last two rather than the first agreeing pair, because a noisy series can match by accident early while still drifting, and what matters is that the run had settled by the time it ended. Otherwise the result is marked NOT CONVERGED; if you see that flag, run more chunks rather than reporting the number. A 30-day run at --chunk-days 5 gives 6 chunks, which is comfortably enough; 10 days gives 2 and will correctly refuse.

GPU utilisation is not a convergence signal. XLA's autotuner keeps the device at 95%+ while it is still choosing kernels. "The GPU is pegged" is not evidence a chunk time is steady-state. Only chunk-to-chunk agreement is.

Run on a genuinely free GPU. Always. No exceptions.

This is the hardest rule here. Simulations should always run on free cards; benchmarks especially, because a contended card does not fail loudly — it returns a plausible-looking number that is simply wrong, and you will not be able to tell from the report. Every other precaution in this document is wasted if the card was shared.

On a shared, unscheduled box, verify the target card is genuinely idle — both no compute apps and near-zero resident memory, since either signal alone looks idle for a parked job:

bash
python tools/gpu_util.py          # every GPU and its tenants
python tools/gpu_util.py --free   # free indices; exit 1 if none

tools/benchmark.py calls the same check as a hard pre-flight gate and refuses to start on a busy card. On a scheduled machine (Derecho) this is the scheduler's job instead — see derecho-jcm-runs.

Never stack two runs on one card. It invalidates both timings and can OOM the other tenant.

Do not start other GPU work anywhere on the box mid-A/B. This is the non-obvious one. Even on a different, genuinely free card, a concurrent job competes for host CPU, PCIe bandwidth and the allocator. If it overlaps only the candidate arm and not the baseline arm, it contaminates exactly one side of the comparison — an asymmetry that shows up as a fake speedup or regression and is invisible in the report. Halving your wall time is not worth a result you then cannot trust. Let the whole interleaved sequence finish on one card before starting anything else.

The corollary: parallelising an A/B across cards is not a shortcut. If you must (because a card is going away), run both arms on card A and both on card B, and report the pairs separately — never baseline on A against candidate on B.

Compare like with like on chunk size. Every chunk boundary costs a host sync, a health check and a netCDF write. The short benchmark uses 5-day chunks (to get ~6 chunks and so detect convergence) while a 12-month run uses 30-day chunks, so the short number carries slightly more per-day write overhead than production. That is a couple of percent, and it cancels entirely in an A/B at the same --chunk-days — but do not quote a 5-day-chunk number as the production throughput.

The run must be a real one. This is a production configuration made shorter and more instrumented, not a reduced-physics proxy: real orography, real SSTs, the packaged CMIP6 ozone, JW init and the production sponge. Check the log line forcing.ozone_file=auto resolved to .../t63/ozone.nc. On a hybrid grid an unresolvable auto now raises rather than degrading, so this cannot pass silently; on a sigma grid it still warns about the ANALYTIC profile, and there the radiation is seeing ~7.6x the tropospheric ozone column and the benchmark is measuring the wrong workload.

A NaN'd run is not a benchmark. The tool reports nan_any and exits non-zero. Timing from a run that blew up mid-flight is meaningless — fix the configuration and re-run rather than quoting the chunks before the blow-up.

Configuration

The presets are not physics=X grid=Y — they carry the whole known-stable override set, because a T63L47 run from an isothermal cold start with no sponge goes NaN within days, which silently destroys the benchmark. Each T63 preset pins init=jw init.rh=0.0, real terrain and forcing from file, and run=longrun — which carries ECHAM's upper sponge. Ozone comes from ozone_file: auto, the shipped default, which resolves the packaged CMIP6 climatology; the preset deliberately does not override it.

Adding a preset: put the complete validated override set in PRESETS, and verify it completes the short benchmark NaN-free before using it for comparisons.

--save-interval is clamped to --chunk-days by the tool: a chunk with zero output times dies in to_xarray() with an opaque IndexError: index 0 is out of bounds for axis 0 with size 0.

Show full SKILL.md (815 more words)Show less

A/B'ing a library version

jcm, jax-rrtmgp and mam4-jax are editable installs — the working tree is the running code. Never git checkout a different branch in the shared clone to benchmark it; that silently changes the code under anyone else's concurrent run.

Instead make a worktree and point the benchmark at it:

bash
git -C /data/dwatsonparris/jax-rrtmgp worktree add \
    /data/dwatsonparris/jax-rrtmgp-worktrees/perf origin/<branch>
$PY tools/benchmark.py ... --pythonpath /data/dwatsonparris/jax-rrtmgp-worktrees/perf

Verify the override actually took before trusting the result:

bash
PYTHONPATH=<worktree> python -c "import rrtmgp,os; print(os.path.dirname(rrtmgp.__file__))"

Run baseline and candidate on the same GPU model, ideally the same card, and interleave if the box is noisy. Report both absolute numbers, not just the ratio.

Pin EVERY editable install, not just the one under test. Pinning only the library you are varying leaves the others free to move mid-sweep — and they do: a six-config sweep here had jax-rrtmgp switch from a feature branch to main between config 1 and config 2 because someone merged a PR, so the two were measuring different radiation code. Nothing failed; the numbers were simply incomparable. It was caught only because the report records resolved SHAs. Put all of jcm, jax-rrtmgp and mam4-jax on --pythonpath as pinned worktrees for any run whose numbers you intend to compare across hours.

Pre-flight by BUILDING the model, not by composing the config. --cfg job resolves the config and stops; it never calls build_model(), so backend-specific rejections pass it and then fail instantly on the GPU. Two have bitten here: an invalid spectral truncation (coords are not built by --cfg job), and init=jw on the pySES backend, which rejects it because it initialises from its own resting USSA-1976 state.

bash
JAX_PLATFORMS=cpu python -c "
from hydra import compose, initialize_config_dir
from jcm.runners import build_model
import os
with initialize_config_dir(config_dir=os.path.abspath('jcm/config'), version_base=None):
    cfg = compose(config_name='config', overrides=[...])
build_model(cfg)"

Do not port overrides between backends by analogy. The pySES preset needed fewer overrides than the spectral one, not the same set: it ignores grid and run.time_step, does its own tracer sub-cycling instead of semi-Lagrangian advection, interpolates the packaged boundary fields onto its columns, and rejects init=jw. Carrying that last one over "for consistency" is exactly how it got in.

Check the precision the configuration needs before launching. f32 (MAM4_JAX_ENABLE_X64=0) is required above T63 and is forward-only. It changes both memory and speed, so it cannot be switched on partway through a sweep without making the halves incomparable — decide once, for all configs. Memory scales with both axes: T63L47 18 GiB, T63L95 61 GiB, T106L47 62 GiB on an 80 GiB card, so T106L95 needs both scalings at once and will not fit in f64.

Benchmark A/Bs must attribute to a single change. A comparison across a merge commit that bundled two independent changes is not evidence about either one — that error is precisely what made the jax-rrtmgp#22 diagnosis wrong (the clamps were blamed; the real cost was a while_loop→scan rewrite in the same PR). If the delta spans more than one change, bisect before attributing.

Interpreting the report

  • s_per_sim_day — the primary comparable number.
  • sim_years_per_day — for provisioning ("can we afford a 100-year run?").
  • peak_mem_gib — includes the compile-time allocation spike, which is a real provisioning requirement. f32 is required above T63 on 80 GB cards.
  • median_util_pct / median_power_w — computed over post-compile samples only. Power well below the card's rating alongside high utilisation indicates a memory-bound or launch-bound workload, not a compute-bound one.

Known reference points

Re-measure rather than trusting these; they are here to catch order-of- magnitude mistakes.

  • Radiation is ~87% of an ECHAM+RRTMGP step, so any RRTMGP change dominates.
  • The AeroCom diagnostic groups cost ~9.4% together at T63L47.
  • f32 (MAM4_JAX_ENABLE_X64=0) is required above T63 and is forward-only: MAM4 microphysics gradients are non-finite in f32.
MA resolution sweep — A100-80GB, 2026-08-11

physics=echam-jam + 2M + semi-Lagrangian, f32, 30 days, jax-rrtmgp at 848da33 (minor-gas scan fix). Grid-native boundary data. Every entry converged, completed 30/30 days, zero NaN.

gridlevelssim days/hrs/sim-daypeak GiButil
T6347127.128.316.688 %
T639565.854.732.692 %
T1064757.462.732.793 %
T1069527.6130.660.296 %
ne30 (pySES)477.2500.161.7100 %
ne30 (pySES)95——OOM at 62—

Not cluster-specific — a vanilla A100-80GB, reproduced within 1–4 % across two machines and both the PCIe and SXM4 variants. A 40 GB card (Derecho) runs the same speeds where the config fits, but on this table only the top two rows do.

Reading it:

  • Levels cost linearly (1.93–2.08× for 2.02× the levels).
  • Resolution is sublinear — T106 has 2.78× T63's columns for 2.21–2.38× the cost, as utilisation climbs 88 % → 96 %. T63L47 leaves the card partly idle; the larger spectral configs are the efficient ones.
  • T106L95 at 60.2 GiB is the largest spectral config that fits 80 GB.
  • pySES ne30 is in a different regime — far more expensive per column and per cell than any spectral config, and radiation is not what dominates it, so the ~87 % rule above does not transfer. See jax-gcm#595; do not extrapolate between the two backends in either direction.

Chunk length does not affect the measurement. ne30L47 at chunk_days 5 and 1 gave 500.07 and 499.7 s/sim-day — 0.07 % apart, all 30 chunk-1 walls inside 498.3–499.7 s. A config re-chunked to fit memory stays comparable.

© climate-analytics-lab, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .claude/skills/jcm-benchmark of climate-analytics-lab/jax-gcm.

Open the folder on GitHubat commit 20ca89d

Compare with similar skills

Jcm Benchmark next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Jcm Benchmark compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Jcm Benchmark this skillclimate-analytics-lab/jax-gcm108—~3.7kAutomated safety check: PassApache-2.0
Add Uint Supportpytorch/pytorch104k2 repos~2.3kAutomated safety check: PassCustom licence
Segment Anything Model GuideOrchestra-Research/AI-Research-SKILLs13k9 repos~3.3kAutomated safety check: PassMIT
Add Oponnx/onnx22k—~1.2kAutomated safety check: PassApache-2.0
CLIP Image-Text MatchingOrchestra-Research/AI-Research-SKILLs13k8 repos~1.7kAutomated safety check: PassMIT
Add Function Bodyonnx/onnx22k—~1.1kAutomated safety check: PassApache-2.0

Similar skills

  • Add Uint Support

    pytorch/pytorch

    Add unsigned integer (uint) type support to PyTorch operators by updating ATDISPATCH macros.

    104k GitHub starsUsed in 2 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Segment Anything Model Guide

    Orchestra-Research/AI-Research-SKILLs

    Guide to using Meta's Segment Anything Model for zero-shot image segmentation with point, box or mask prompts, or automatic mask generation.

    13k GitHub starsUsed in 9 repos~3.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Add Op

    onnx/onnx

    Add a new ONNX operator or update an existing operator to a new opset version.

    22k GitHub stars~1.2k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • CLIP Image-Text Matching

    Orchestra-Research/AI-Research-SKILLs

    Explains OpenAI's CLIP model for zero-shot image classification, image-text similarity, semantic image search and content moderation, with install steps and code patterns.

    13k GitHub starsUsed in 8 repos~1.7k tokens
    AI & LLM EngineeringAuto-check passed
  • Add a function body definition to an ONNX operator, defining how it decomposes into simpler ops.

    22k GitHub stars~1.1k tokensUpdated yesterday
    AI & LLM EngineeringAuto-check passed
  • Paddle Design Distributed

    PaddlePaddle/Paddle

    A skill your agent uses when working with Paddle's distributed training system: understanding parallelism strategies (DP, ZeRO, TP, PP, SP), semi-automatic parallel with ProcessMesh + shardtensor…

    24k GitHub stars~660 tokensUpdated 8 days ago
    AI & LLM EngineeringAuto-check passed

More from climate-analytics-lab/jax-gcm

  • Derecho Jcm Runs

    climate-analytics-lab/jax-gcm

    Submit, monitor and benchmark jax-gcm (jcm) simulations on NCAR Derecho's PBS queues.

    108 GitHub stars~2.5k tokensUpdated today
    Auto-check passed
  • Kubernetes Jcm Runs

    climate-analytics-lab/jax-gcm

    Run jcm on a Kubernetes GPU cluster — generate benchmark and production Job manifests, pick a comparable GPU, survive eviction, collect results.

    108 GitHub stars~3.4k tokensUpdated today
    Auto-check passed
  • Devbox Jcm Runs

    climate-analytics-lab/jax-gcm

    Run jcm on the shared UCSD dev workstation (8x A100-80GB, no scheduler) — find a genuinely free GPU, avoid stomping on colleagues' jobs, environment and scratch paths, and the etiquette/traps…

    108 GitHub stars~1.7k tokensUpdated today
    Auto-check passed
  • Jcm Dev Workflow

    climate-analytics-lab/jax-gcm

    End-to-end development workflow for jcm — atomic commits, the local test/lint gate, opening a PR linked to its issue, monitoring CI and the automatic Codex review, addressing feedback, and handing…

    108 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Jcm Local CI

    climate-analytics-lab/jax-gcm

    Run the jax-gcm CI gates locally on Derecho when GitHub Actions minutes are exhausted or a pre-push check is wanted — lint, fast tests (90% coverage), slow tests (80% PR coverage) and a local Claude…

    108 GitHub stars~3k tokensUpdated today
    Auto-check passed
  • Jcm Run

    climate-analytics-lab/jax-gcm

    Launch a jcm model run through the built-in Hydra configs — config groups, the validated stable T63L47 overrides, Hydra override traps, and watching for startup failures.

    108 GitHub stars~2.1k tokensUpdated today
    Auto-check passed

Questions about Jcm Benchmark

What does Jcm Benchmark do?

Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion. Jcm Benchmark is an agent skill from climate-analytics-lab/jax-gcm. Measure jcm throughput reproducibly — short (1 month) or long (12 month) runs on a validated stable config, with GPU memory/utilisation logging and an explicit convergence criterion.

When should I use Jcm Benchmark?

Jcm Benchmark fits situations like: comparing hardware; library versions (jax-rrtmgp; physics presets; checking for a performance regression.

How do I install Jcm Benchmark in Claude Code?

Run `npx skills add climate-analytics-lab/jax-gcm --skill jcm-benchmark -a claude-code`. Or copy the skill folder (.claude/skills/jcm-benchmark in climate-analytics-lab/jax-gcm) into .claude/skills/jcm-benchmark in your project. Claude Code loads it when a task matches its description.

How do I install Jcm Benchmark in Codex?

Run `npx skills add climate-analytics-lab/jax-gcm --skill jcm-benchmark -a codex`. Or copy the skill folder (.claude/skills/jcm-benchmark in climate-analytics-lab/jax-gcm) into .agents/skills/jcm-benchmark in your project. Codex loads it when a task matches its description.

Can I use Jcm Benchmark in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add climate-analytics-lab/jax-gcm --skill jcm-benchmark -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/jcm-benchmark, .gemini/skills/jcm-benchmark, .github/skills/jcm-benchmark and .opencode/skills/jcm-benchmark in your project.

What does Jcm Benchmark need to run?

Going by SKILL.md and its folder, Jcm Benchmark needs the command-line tools its instructions call (python and git). Our summary lists: Python 3.

Does Jcm Benchmark access the network?

SKILL.md contains no URLs. Its commands use git, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Jcm Benchmark safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Jcm Benchmark use?

Jcm Benchmark is published under the Apache-2.0 licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Jcm Benchmark use?

About 3.7k tokens (SKILL.md is roughly 15k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Jcm Benchmark?

Skills that share tags, products or a category with Jcm Benchmark: Add Uint Support (pytorch/pytorch, 104k stars), Segment Anything Model Guide (Orchestra-Research/AI-Research-SKILLs, 13k stars), Add Op (onnx/onnx, 22k stars) and CLIP Image-Text Matching (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Jcm Benchmark?

climate-analytics-lab (a GitHub organization) maintains it in climate-analytics-lab/jax-gcm, which has 108 GitHub stars. The repository holds 7 skills in this directory. The repository was last updated on October 7, 2026.

Source: climate-analytics-lab/jax-gcm on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.