Agent skill

Weather Data Reproducibility

by sickn33 in sickn33/agentic-awesome-skills

Record and verify provenance manifests for weather-data inputs and derived artifacts, including object identity, selections, software versions, transformations, and hashes.

MITAuto-check passedResearch & Science

Install Weather Data Reproducibility

skills CLI
$ npx skills add sickn33/agentic-awesome-skills --skill weather-data-reproducibility -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install sickn33/agentic-awesome-skills weather-data-reproducibility --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/sickn33/agentic-awesome-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/weather-data-reproducibility .claude/skills/weather-data-reproducibility && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
weather-data-reproducibility
GitHub stars
47k
Used in
1 other repo
Token cost
~2.9k tokens
SKILL.md length
1,232 words
Files
1
Skills in repo
1,497
Repo updated
First seen
Licence
MIT

At a glance

Record and verify provenance manifests for weather-data inputs and derived artifacts, including object identity, selections, software versions, transformations, and hashes.

  • Works in 8 steps: Validate the manifest schema version and… → Re-resolve every immutable object… → Download or locate inputs and compare… → …
  • Tasks that involve Physical and earth sciences
  • SKILL.md covers Overview, When to Use This Skill, Reproducibility Contract and Minimum Manifest, plus 12 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Weather Data Reproducibility is an agent skill from sickn33/agentic-awesome-skills. Record and verify provenance manifests for weather-data inputs and derived artifacts, including object identity, selections, software versions, transformations, and hashes.

Its SKILL.md is about 2.9k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Research & Science, covering Physical and earth sciences and Reproducible research. The repository describes itself as: AAS Core is the local, agent-first control plane for complete catalog discovery, agent-owned selection, stack validation, and planning, backed by 2,400+ agentic skills. Includes… The licence is MIT.

When your agent uses it

  • Tasks that involve Physical and earth sciences
  • Tasks that involve Reproducible research

Example prompts

  • “/weather-data-reproducibility”

Requirements

  • Python 3

Workflow steps

8 steps, taken from the first numbered list in SKILL.md.

  1. Validate the manifest schema version and required fields.
  2. Re-resolve every immutable object identity. Fail if an object changed unless
  3. Download or locate inputs and compare size plus a real checksum.
  4. Reapply the recorded selection and confirm the materialized subset hash.
  5. Recreate the pinned processing environment and transformations.
  6. Produce artifacts in a new output directory.
  7. Compare artifact hashes for bitwise claims or documented scientific metrics
  8. Write a new run manifest that references, but does not overwrite, the

What it can do on your machine

Read from SKILL.md and the folder at commit b84d35a. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are json and python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Links to these hosts (documentation or services it may open):

    • docs.aws.amazon.com
    • w3.org
    • cfconventions.org
    • ncei.noaa.gov

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Weather Data Reproducibility loads about 2.9k tokens when it runs. Until then it costs about 50 tokens; SKILL.md has 1,232 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~50
When it runs · the whole SKILL.md, loaded when a task matches
~2.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from sickn33/agentic-awesome-skills at commit b84d35a, republished under its MIT licence (© sickn33). 1,232 words, ~2,901 tokens.

Download SKILL.mdSave it as .claude/skills/weather-data-reproducibility/SKILL.md (or your agent's skills folder).
name
weather-data-reproducibility
description
Record and verify provenance manifests for weather-data inputs and derived artifacts, including object identity, selections, software versions, transformations, and hashes.
category
data
risk
safe
source
self
source_type
self
date_added
2026-09-24
author
ShianMike
tags
weather, reproducibility, provenance, checksums, grib2, netcdf
tools
claude, cursor, gemini, codex

Weather Data Reproducibility

Overview

Create a compact provenance manifest that identifies the exact weather inputs, subsets, software, transformations, and output artifacts used by a run. Verify the manifest before claiming that a result can be replayed or reproduced.

This skill records and checks lineage. It does not discover or download data by itself; pair it with the appropriate model, observation, radar, or satellite fetching skill.

When to Use This Skill

  • A map, sounding, time series, dataset, or benchmark must be auditable.
  • Weather inputs came from mutable URLs, mirrors, byte-range subsets, or caches.
  • A result must distinguish model initialization, forecast lead, and valid time.
  • Temporary GRIB2, NetCDF, radar, satellite, or observation files will be deleted after processing.
  • Two outputs differ and their inputs or processing environment must be compared.
  • A publication or release needs machine-readable data and software provenance.

Do not claim bit-for-bit reproducibility merely because a source URL and run time were logged.

Reproducibility Contract

Decide which claim the workflow supports:

  • Traceable: The source and processing history can be audited.
  • Replayable: The same identified inputs can still be retrieved and the recorded procedure can be rerun.
  • Bitwise reproducible: The rerun is expected to produce identical bytes in a pinned environment.
  • Scientifically reproducible: Equivalent inputs and methods produce the same scientific conclusion within a stated tolerance.

Record the intended claim and its tolerance. Mutable remote objects, floating software versions, lossy images, parallel reductions, and platform-dependent code can make the stronger claims impossible.

Minimum Manifest

Use a versioned JSON document. Keep request intent separate from the source that was actually resolved.

json
{
  "schema_version": 1,
  "created_utc": "2026-09-18T09:00:00Z",
  "claim": {"level": "replayable", "tolerance": null},
  "request": {
    "dataset": "hrrr",
    "product": "prs",
    "requested_valid_time": "2026-09-18T12:00:00Z",
    "variables": ["temperature", "relative_humidity"]
  },
  "resolved": {
    "initialization": "2026-09-18T06:00:00Z",
    "forecast_hour": 6,
    "valid_time": "2026-09-18T12:00:00Z",
    "provider": "aws"
  },
  "inputs": [
    {
      "role": "pressure_fields",
      "uri": "s3://bucket/exact-object-key",
      "object_identity": {
        "etag": "opaque-identity",
        "last_modified": "2026-09-18T07:01:02Z",
        "size_bytes": 987654321
      },
      "selection": {
        "inventory_uri": "s3://bucket/exact-object-key.idx",
        "inventory_sha256": "HEX_DIGEST",
        "byte_ranges": [[1000, 1999], [5000, 6999]]
      },
      "materialized_sha256": "HEX_DIGEST"
    }
  ],
  "processing": {
    "software": {"application": "1.2.3", "python": "3.13.7"},
    "parameters": {"point": [35.22, -97.44], "nearest_method": "model-aware"},
    "transformations": ["decode GRIB2", "convert K to degC"]
  },
  "artifacts": [
    {
      "path": "sounding.png",
      "media_type": "image/png",
      "size_bytes": 123456,
      "sha256": "HEX_DIGEST"
    }
  ]
}

Use null for an intentionally absent value and omit fields whose meaning is unknown. Never fill a required-looking field with a guess.

Identify Every Input

For each remote object or API response, record as available:

  • provider, dataset, endpoint, bucket, and exact key or immutable URL;
  • version ID, generation number, release, or dataset revision;
  • ETag as opaque object identity, Last-Modified, and content length;
  • provider-supplied checksum and checksum algorithm;
  • a locally computed SHA-256 for downloaded or materialized bytes;
  • response media type and content encoding;
  • retrieval time and successful fallback provider;
  • license or attribution identifier when required.

An S3 ETag is not always an MD5 digest, particularly for multipart or encrypted objects. Store it for identity and compute a cryptographic content hash when the actual bytes must be verified.

Record the Selection

The source object alone is insufficient when only part of it was used. Record:

  • GRIB2 inventory URI and inventory content hash;
  • exact inventory rows, search expression, and inclusive byte ranges;
  • Zarr group, array names, chunk keys, and coordinate slices;
  • API query parameters and a hash or retained copy of the raw response;
  • station identifier scheme and observation time interval;
  • radar site/product and scan interval;
  • satellite platform, product, sector, channel, and scan interval;
  • requested coordinates plus the actual selected grid point and distance.

Preserve the order in which selected binary ranges were assembled. Hash the materialized subset separately from the full remote object's identity.

Record Weather Time Semantics

Use explicit UTC timestamps and name their meaning:

  • models: initialization, forecast lead, and valid time;
  • station observations: observation, correction, and ingestion time;
  • radar: volume or product scan start and end;
  • satellite: scan start, scan end, and file creation time;
  • retrieval: when the client fetched the metadata and bytes.

Do not replace these with one ambiguous timestamp field. Record the calendar and leap-second handling if the source or application requires it.

Record Processing and Environment

Capture only details that can change the result:

  • application version and source commit, including dirty-tree status;
  • decoder and scientific library versions;
  • runtime and operating-system or architecture details when relevant;
  • command arguments or structured parameters;
  • unit conversions, QC decisions, interpolation, coordinate selection, and aggregation rules;
  • deterministic random seed when randomness is present;
  • container image digest or environment lock-file hash when available.

Do not dump an entire environment full of unrelated packages merely because it is easy. Prefer a lock file plus the versions of software that actually touched the data.

Hash and Write Atomically

The Python standard library is sufficient for local artifact hashes and a canonical, atomically replaced JSON manifest:

python
import hashlib
import json
import os
from pathlib import Path


def sha256_file(path, chunk_size=1024 * 1024):
    digest = hashlib.sha256()
    with Path(path).open("rb") as handle:
        for chunk in iter(lambda: handle.read(chunk_size), b""):
            digest.update(chunk)
    return digest.hexdigest()


def write_manifest(path, manifest):
    path = Path(path)
    temporary = path.with_suffix(path.suffix + ".tmp")
    payload = json.dumps(
        manifest, ensure_ascii=False, sort_keys=True, separators=(",", ":")
    ) + "\n"
    temporary.write_text(payload, encoding="utf-8", newline="\n")
    os.replace(temporary, path)

Hash artifacts only after their writers are closed. If the manifest itself must be signed, sign the canonical bytes using the project's established signing workflow rather than inventing a custom scheme.

Show full SKILL.md (522 more words)Show less

Verify and Replay

Before replay:

  1. Validate the manifest schema version and required fields.
  2. Re-resolve every immutable object identity. Fail if an object changed unless the reproducibility contract explicitly permits an equivalent replacement.
  3. Download or locate inputs and compare size plus a real checksum.
  4. Reapply the recorded selection and confirm the materialized subset hash.
  5. Recreate the pinned processing environment and transformations.
  6. Produce artifacts in a new output directory.
  7. Compare artifact hashes for bitwise claims or documented scientific metrics and tolerances for scientific claims.
  8. Write a new run manifest that references, but does not overwrite, the original.

Report the first mismatch with its role, expected value, and actual value.

Temporary Data Lifecycle

  • Write the final artifact and its manifest before deleting request-owned input files.
  • Keep cleanup in finally so failures and cancellation do not leak large GRIB2, NetCDF, radar, or observation files.
  • Never delete a shared user cache as request cleanup.
  • If raw inputs are deleted and the remote source is mutable or short-lived, label the run traceable rather than replayable.
  • A manifest is small and should normally outlive transient downloads.

Verification Checklist

  • Request intent and resolved data identity are separate.
  • Every timestamp is UTC and has explicit semantics.
  • Every input has an exact locator, identity metadata, and content hash when bytes were materialized.
  • Subset rules, inventory identity, byte ranges, and selected coordinates are recorded.
  • Software versions and result-changing transformations are explicit.
  • Final artifacts have size, media type, and SHA-256.
  • The claimed reproducibility level matches what can actually be replayed.
  • The manifest contains no secrets or machine-specific private information.

Security & Safety Notes

  • Never record credentials, authorization headers, cookies, signed-query strings, private bucket names, or secret environment variables.
  • Redact user names and unnecessary absolute local paths before sharing a manifest.
  • Treat a hash as an integrity value, not proof that an untrusted input is safe.
  • Validate manifest paths before opening files; do not allow .. traversal or writes outside the intended output directory.
  • Keep TLS verification enabled when re-fetching inputs.
  • Preserve license, attribution, redistribution, and access restrictions.

Common Pitfalls

  • ETag was labeled SHA-256 or MD5: ETag semantics depend on the provider and upload method. Store it as identity and compute a real local hash.
  • A model result cannot be recreated: Initialization and forecast lead were collapsed into valid time. Record all three.
  • The same URI returns different bytes: The remote object was mutable and no version or checksum was pinned. Downgrade the claim or retain the input.
  • A GRIB subset differs: The inventory version, selected rows, or inclusive byte ranges were omitted.
  • The artifact hash changes across systems: The process is scientifically, but not bitwise, reproducible. Define and test a numerical tolerance.

Limitations

  • A manifest cannot restore data that was deleted from a mutable source unless the input bytes or an immutable archive were retained.
  • Exact replay may require licensed software, hardware, or provider access not captured in a public manifest.
  • Provenance establishes lineage and integrity; it does not establish forecast skill, observational truth, or scientific validity.

Additional Resources

© sickn33, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in skills/weather-data-reproducibility of sickn33/agentic-awesome-skills.

Open the folder on GitHubat commit b84d35a

Used in 1 other repository

We found 5 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 1 other GitHub owner. This page covers the copy in sickn33/agentic-awesome-skills, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Weather Data Reproducibility next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Weather Data Reproducibility compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Weather Data Reproducibility this skillsickn33/agentic-awesome-skills47k1 repos~2.9kAutomated safety check: PassMIT
Lammps Evidence MdCai-aa/CAE-Agent-Hub1k—~476Automated safety check: PassMIT
Mechanical Engineering Researchhashgraph-online/awesome-codex-plugins1.3k—~2.9kAutomated safety check: PassApache-2.0
Ngeo Methodsbrycewang-stanford/Awesome-Journal-Skills1.2k—~1.5kAutomated safety check: PassMIT
Peer ReviewK-Dense-AI/claude-scientific-writer2.4k2 repos~3.1kAutomated safety check: NotesMIT
AstropyzLanqing/codex-claude-academic-skills4.7k13 repos~2.9kAutomated safety check: PassBSD-3-Clause

Similar skills

  • Lammps Evidence Md

    Cai-aa/CAE-Agent-Hub

    Evidence-first LAMMPS molecular-dynamics workflow for input decks, potentials, minimization, equilibration, deformation, thermo logs, trajectories, restart/data files, stress-strain extraction, and…

    1k GitHub stars~476 tokensUpdated 10 days ago
    Research & ScienceAuto-check passed
  • Mechanical Engineering Research

    hashgraph-online/awesome-codex-plugins

    Apply source-aware mechanical-engineering judgment to research, analysis, coding, writing, teaching, research identity, and release work.

    1.3k GitHub stars~2.9k tokensUpdated yesterday
    Research & ScienceAuto-check passed
  • Ngeo Methods

    brycewang-stanford/Awesome-Journal-Skills

    A skill your agent uses when building the online Methods section of a Nature Geoscience manuscript so every quantitative Earth-science claim is grounded in data, model diagnostics, and quantified…

    1.2k GitHub stars~1.5k tokensUpdated 13 days ago
    Research & ScienceAuto-check passed
  • Peer Review

    K-Dense-AI/claude-scientific-writer

    Prepare evidence-bounded, constructive peer-review drafts and structured manuscript assessments.

    2.4k GitHub starsUsed in 2 repos~3.1k tokens
    Research & ScienceAuto-check: notes
  • Astropy

    zLanqing/codex-claude-academic-skills

    Comprehensive Python library for astronomy and astrophysics.

    4.7k GitHub starsUsed in 13 repos~2.9k tokens
    Research & ScienceAuto-check passed
  • Guides an agent through reproducing papers built on the CHARLS health and retirement survey, from variable mapping to cognition, depression and isolation scores.

    617 GitHub starsUsed in 1 repo~1.8k tokens
    Research & ScienceAuto-check passed

More from sickn33/agentic-awesome-skills

All 1,497 skills in this repo
  • Liuguang Banlan UI

    sickn33/agentic-awesome-skills

    Implements an interface in one of two named color modes, iridescent white or colorful black, from a parameterized starter that reports measured color intensity.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • User Thoughts Memory

    sickn33/agentic-awesome-skills

    Saves a user's project decisions, rules and preferences into a project-local mdbase so later sessions and other agents can recover the intent.

    47k GitHub starsUsed in 1 repo~2.5k tokens
    Auto-check passed
  • Using LWC Memory and Graphs

    sickn33/agentic-awesome-skills

    Keeps project decisions, research and verified results available across coding-agent sessions through LWC memory, a document Wiki graph and a CodeGraph code index.

    47k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Find Complementary Founders

    sickn33/agentic-awesome-skills

    Guides an agent through assessing its own owner for cofounder fit, publishing an approved profile, and ranking complementary profiles other agents published for their owners.

    47k GitHub starsUsed in 1 repo~4.8k tokens
    Auto-check passed
  • Whatsapp Cloud API

    sickn33/agentic-awesome-skills

    Integracao com WhatsApp Business Cloud API (Meta). An agent skill from sickn33/agentic-awesome-skills.

    47k GitHub starsUsed in 2 repos~4.5k tokens
    Auto-check passed
  • Cline Pilot

    sickn33/agentic-awesome-skills

    Acts as a proxy for the Cline CLI, dispatching coding tasks one at a time, monitoring runs by hard evidence, relaying decisions to you and learning per-project preferences.

    47k GitHub starsUsed in 1 repo~4.6k tokens
    Auto-check passed

Questions about Weather Data Reproducibility

What does Weather Data Reproducibility do?

Record and verify provenance manifests for weather-data inputs and derived artifacts, including object identity, selections, software versions, transformations, and hashes. Weather Data Reproducibility is an agent skill from sickn33/agentic-awesome-skills. Record and verify provenance manifests for weather-data inputs and derived artifacts, including object identity, selections, software versions, transformations, and hashes.

When should I use Weather Data Reproducibility?

Weather Data Reproducibility fits situations like: tasks that involve Physical and earth sciences; tasks that involve Reproducible research.

How do I install Weather Data Reproducibility in Claude Code?

Run `npx skills add sickn33/agentic-awesome-skills --skill weather-data-reproducibility -a claude-code`. Or copy the skill folder (skills/weather-data-reproducibility in sickn33/agentic-awesome-skills) into .claude/skills/weather-data-reproducibility in your project. Claude Code loads it when a task matches its description.

How do I install Weather Data Reproducibility in Codex?

Run `npx skills add sickn33/agentic-awesome-skills --skill weather-data-reproducibility -a codex`. Or copy the skill folder (skills/weather-data-reproducibility in sickn33/agentic-awesome-skills) into .agents/skills/weather-data-reproducibility in your project. Codex loads it when a task matches its description.

Can I use Weather Data Reproducibility in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add sickn33/agentic-awesome-skills --skill weather-data-reproducibility -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/weather-data-reproducibility, .gemini/skills/weather-data-reproducibility, .github/skills/weather-data-reproducibility and .opencode/skills/weather-data-reproducibility in your project.

What does Weather Data Reproducibility need to run?

SKILL.md names no scripts, command-line tools or credentials: Weather Data Reproducibility is instructions for the agent only. Our summary lists: Python 3.

Does Weather Data Reproducibility access the network?

SKILL.md names 4 domains. As links in the text: docs.aws.amazon.com, w3.org, cfconventions.org and ncei.noaa.gov. This is read from the text; nothing was executed.

Is Weather Data Reproducibility safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Weather Data Reproducibility use?

Weather Data Reproducibility is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Weather Data Reproducibility use?

About 2.9k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Weather Data Reproducibility?

Skills that share tags, products or a category with Weather Data Reproducibility: Lammps Evidence Md (Cai-aa/CAE-Agent-Hub, 1k stars), Mechanical Engineering Research (hashgraph-online/awesome-codex-plugins, 1.3k stars), Ngeo Methods (brycewang-stanford/Awesome-Journal-Skills, 1.2k stars) and Peer Review (K-Dense-AI/claude-scientific-writer, 2.4k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Weather Data Reproducibility?

sickn33 (a GitHub user) maintains it in sickn33/agentic-awesome-skills, which has 47,405 GitHub stars. The repository holds 1,497 skills in this directory. The repository was last updated on October 9, 2026.

Source: sickn33/agentic-awesome-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.