Official agent skill

Earth2studio Create Datasource

by NVIDIA in NVIDIA/skills

Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores.

OfficialApache-2.0Auto-check passedDevelopment

Install Earth2studio Create Datasource

skills CLI
$ npx skills add NVIDIA/skills --skill earth2studio-create-datasource -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/skills earth2studio-create-datasource --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/earth2studio-create-datasource .claude/skills/earth2studio-create-datasource && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
earth2studio-create-datasource
GitHub stars
3.5k
Token cost
~2.6k tokens
SKILL.md length
1,083 words
Files
19 (incl. references)
Skills in repo
380
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores.

  • Works in 12 steps: Obtain Remote Data Store Reference → Determine Source Type → Examine Remote Store & Propose… → …
  • Fetching data with existing sources
  • SKILL.md covers Purpose, Prerequisites, Workspace and Instructions, plus 3 more sections
  • Runs Python and Shell scripts from its folder; calls uv, make and gh

What it does

Earth2studio Create Datasource is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores. Do NOT use for fetching data with existing sources, model inference, or installation tasks.

Its SKILL.md is about 2.6k tokens, which your agent loads only when the skill is triggered. The skill folder holds 24 other files, including reference files (for example `BENCHMARK.md`, `evals/config.yml` and `evals/environment/setup/bootstrap.sh`).

It sits in Development. It works with Python. The repository describes itself as: Agent Skills for NVIDIA products — install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. The licence is Apache-2.0.

When your agent uses it

  • Fetching data with existing sources
  • Model inference
  • Installation tasks

Example prompts

  • “/earth2studio-create-datasource”

Requirements

  • Python 3
  • A Bash shell

Workflow steps

12 steps, taken from the step headings in SKILL.md.

  1. Obtain Remote Data Store Reference
  2. Determine Source Type
  3. Examine Remote Store & Propose Dependencies
  4. Add Dependencies
  5. Create Lexicon Class
  6. Update E2STUDIO_VOCAB / SCHEMA (if needed)
  7. Create Skeleton Data Source File
  8. Implement the Data Source & Tests
  9. Register the Source
  10. Update Documentation
  11. Update CHANGELOG.md
  12. Verify Style & Expand Tests

What it can do on your machine

Read from SKILL.md and the folder at commit 0e0d506. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Ships script files (Python and Shell, from the files we listed), which the agent can run.

    Shell commands in SKILL.md call:

    • uv
    • make
    • gh

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv and gh, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Earth2studio Create Datasource loads about 2.6k tokens when it runs, and up to ~12k if it reads all its reference files. Until then it costs about 66 tokens; SKILL.md has 1,083 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~66
When it runs · the whole SKILL.md, loaded when a task matches
~2.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~12k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/skills at commit 0e0d506, republished under its Apache-2.0 licence (© NVIDIA). 1,083 words, ~2,627 tokens.

Download SKILL.mdSave it as .claude/skills/earth2studio-create-datasource/SKILL.md (or your agent's skills folder). This skill also uses 18 other files; get the full folder from GitHub.
name
earth2studio-create-datasource
description
Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores. Do NOT use for fetching data with existing sources, model inference, or installation tasks.
version
0.16.0
license
Apache-2.0
metadata.author
NVIDIA Earth-2 Team <agent-skills@nvidia.com>
metadata.tags
earth2studio, earth2, python, data-source, forecast-source, integration
argument-hint
URL or description of remote data store (optional)

Create and Validate Data Source

Purpose

End-to-end workflow for implementing a new Earth2Studio data source wrapper that connects a remote data store (S3, GCS, Azure, HTTP, HuggingFace) to Earth2Studio's async data fetching infrastructure — from analysis through implementation, testing, validation, and PR submission.

Prerequisites

  • Earth2Studio dev environment with uv (uv run python must work)
  • Git configured with fork (origin) and upstream (upstream) remotes
  • Access to the target remote data store (credentials if private)
  • Python 3.10+

Workspace

Use the directory containing pyproject.toml. For Harbor evals, write to /workspace/output/ preserving paths. Never read evals/targets/.

Instructions

Python Environment: Always use uv run python or the local .venv. Never use the system Python directly.

Follow every step in order.

[CONFIRM] gates: Only Step 1 (Source Type) and Step 12 (Sanity-Check Plots) require explicit user approval. All other [CONFIRM] markers are advisory — present decisions inline and proceed without blocking.

Deliverables first: Write the source file and test file (Steps 6–7) before extended exploration, documentation, registration, CHANGELOG, or PR work. Skip Steps 8–14 when the user asks for implementation only.

Before you finish: Run verification commands in the repo root so results appear in the session log:

bash
uv run pytest test/data/test_<source>.py -x
make format && make lint

Be concise: Avoid long architecture reports; summarize decisions in a few sentences and move on to file writes.

Hangs or User Feedback If agent becomes stuck or user provides a correction during this skills use, conservatively review relevant part of the skill and improve. Be concise.

One source type per invocation. Invoke again for companion types.

Reference Files

Load these on demand during the relevant steps:

FileContentLoad at
references/implementation-guide.pySkeleton source with FILL commentsSteps 3–10
references/testing-guide.pyTest skeleton with FILL commentsStep 11
references/validation-guide.mdPlot templates, PR body template, Greptile handlingSteps 12–14 (optional, for templates)

Workflow Overview
text
Step 0: Obtain reference → Step 1: Determine type → Step 2: Dependencies
→ Step 3: Add deps → Step 4: Create lexicon → Step 5: Update vocab/schema
→ Step 6: Create skeleton → Step 7: Implement source → Step 8: Register
→ Step 9: Documentation → Step 10: CHANGELOG → Step 11: Tests
→ Step 12: Validate & plots (user confirms) → Step 13: PR + sanity comment
→ Step 14: Greptile review

Step 0 — Obtain Remote Data Store Reference

If $ARGUMENTS is provided, use it (URL → WebFetch; file path → read).

If empty, ask:

Please provide a URL, API documentation link, or description of the remote data store. This will be used to understand storage format, access pattern, variable inventory, temporal/spatial resolution.


Step 1 — Determine Source Type
ProtocolReturnsHas lead_time?Use
DataSourcexr.DataArrayNoGridded analysis/reanalysis
ForecastSourcexr.DataArrayYesGridded forecast
DataFrameSourcepd.DataFrameNoSparse/station obs
ForecastFrameSourcepd.DataFrameYesSparse forecast obs

Key factors: gridded vs sparse → DataArray vs DataFrame; analysis vs forecast → Source vs ForecastSource.

[CONFIRM — Source Type]

Present recommended type with justification. Ask for confirmation.


Step 2 — Examine Remote Store & Propose Dependencies

Analyze: storage backend, file format, authentication, access pattern, temporal/spatial resolution, variable inventory.

Prefer fsspec:

BackendPreferredAvoid
AWS S3s3fs (core dep)boto3 directly
GCSgcsfs (core dep)google-cloud-storage
Azureadlfsazure-storage-blob
HTTPfsspec (core dep)requests
HuggingFacehuggingface_hub (core dep)custom scripts

Only fall back to dedicated libraries when fsspec cannot access the store.

Check pyproject.toml — only propose packages not already present. Core deps include: s3fs, gcsfs, fsspec, zarr, netCDF4, h5py, pygrib, huggingface-hub, pandas, pyarrow.

[CONFIRM — Dependencies & Access Pattern]

Present: backend, fsspec filesystem, new packages (with license), auth method.


Step 3 — Add Dependencies

Load references/implementation-guide.py from here through Step 10.

If new packages needed:

  1. uv add --extra data <package>
  2. uv lock
  3. Add optional dependency imports using OptionalDependencyFailure pattern

Step 4 — Create Lexicon Class

Create earth2studio/lexicon/<source_name>.py with:

  • metaclass=LexiconType
  • VOCAB: dict[str, str] mapping E2S names → remote keys
  • get_item(cls, val) returning tuple[str, Callable]
  • Use :: separator for structured keys

Map remote variables against E2STUDIO_VOCAB (282 entries in earth2studio/lexicon/base.py).

[CONFIRM — Lexicon & Variable Mapping]

Present: class name, key format, full mapping table, modifiers, reference URL.


Step 5 — Update E2STUDIO_VOCAB / SCHEMA (if needed)
  • New vocab: surface = descriptive abbrev; pressure = {name}{level}
  • New schema fields: DataFrame sources only, check E2STUDIO_SCHEMA
[CONFIRM — Vocabulary & Schema Updates]

Skip if no updates needed.


Step 6 — Create Skeleton Data Source File

Follow canonical method ordering:

  1. Class constants
  2. SCHEMA
  3. __init__
  4. _async_init
  5. __call__
  6. fetch
  7. _create_tasks
  8. fetch_wrapper
  9. fetch_array
  10. _validate_time
  11. Helpers
  12. cache property
  13. available classmethod

Use async task dataclass pattern for parallel execution.

Show full SKILL.md (424 more words)Show less
[CONFIRM — Skeleton]

Present: class name, file path, skeleton code, task dataclass.


Step 7 — Implement the Data Source & Tests

The test file is a co-equal deliverable. Create test/data/test_<filename>.py alongside the source. Add test_<source>_call_mock for async sources.

Sync sources: Use prep_data_inputs/prep_forecast_inputs, direct __call__.

Async sources: See references/implementation-guide.py for required patterns: _sync_async, managed_session, gather_with_concurrency, async_retry, pure async I/O, try/finally cleanup. Constructor params: cache=True, verbose=True, async_timeout=600, async_workers=16, retries=3. DataFrame sources add time_tolerance.


Step 8 — Register the Source
  • earth2studio/data/__init__.py — alphabetical import
  • earth2studio/lexicon/__init__.py — alphabetical import
  • Verify pyproject.toml deps

Step 9 — Update Documentation
  • Add to correct RST file (datasources_analysis.rst / _forecast.rst / _dataframe.rst)
  • Class docstring: Parameters, Warning (download size), Note (reference URLs), Badges (last)
  • All public methods: NumPy-style docstrings

Step 10 — Update CHANGELOG.md

Add entry under the current unreleased version. See references/implementation-guide.py REGISTRATION CHECKLIST for the format.

One line per source. Do NOT add separate lexicon entries.


Step 11 — Verify Style & Expand Tests

Run make format && make lint && make license. Load references/testing-guide.py for test skeletons. Required tests: test_<source>_fetch (slow), _cache (slow), _call_mock, _exceptions, _available. Target 90%+ coverage with --slow.

[CONFIRM — Tests]

Present test file, functions, coverage.


Step 12 — Validate Variables & Sanity-Check
  1. Validate all lexicon vars against real data (run script, do NOT commit)
  2. Remove variables with < 10% valid data
  3. Create sanity-check plot (gridded or sparse template)
  4. Tell user the plot path and ask for visual confirmation
[CONFIRM — Sanity-Check Plots]

User MUST visually inspect plots. Do not proceed without confirmation.


Step 13 — Branch, Commit & Open PR
  1. Create branch feat/data-source-<name>
  2. Commit (do NOT add sanity-check script/images)
  3. Push to fork
  4. gh pr create --repo NVIDIA/earth2studio
  5. Immediately post sanity-check validation as PR comment with:
    • Variable coverage table (name, count, range, unit)
    • Data validation summary (regions, storms/stations, time range)
    • Key findings (physically reasonable values, conversions verified)
    • Full validation script in <details> block
    • Image placeholder: <!-- Drag and drop sanity-check image here -->
[CONFIRM — Ready to Submit]

Verify all steps complete before creating PR.


Step 14 — Automated Code Review
  1. Poll for Greptile review (5 min timeout)
  2. Categorize feedback (bug/style/perf/docs/suggestion/false-positive)
  3. Present triage table to user
  4. Implement accepted fixes
  5. Respond to PR comments
  6. Push
[CONFIRM — Review Triage]

User approves which comments to address.


Examples

text
User: Add a data source for the NOAA GFS analysis on S3
Agent: [loads skill, proceeds through Steps 0–14]

Limitations

  • One source type per invocation
  • Requires network access for validation (Step 12)
  • SPDX license headers required in all files

Reminders

DO: uv run python, loguru.logger, alphabetical order in __init__.py/RST/CHANGELOG, canonical method ordering, async utilities (managed_session, gather_with_concurrency, async_retry), pure async I/O, reference URLs in docstrings, try/finally cleanup.

AVOID: asyncio.to_thread, bare tqdm.gather, xarray for loading, full file downloads.

NEVER: loop.set_default_executor(), commit secrets, commit sanity-check scripts/images.

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 18 other files (references) in skills/earth2studio-create-datasource of NVIDIA/skills.

  • SKILL.md
  • BENCHMARK.md
  • evals/config.yml
  • evals/environment/Dockerfile
  • evals/environment/setup/bootstrap.sh
  • evals/evals.json
  • evals/targets/eval_1_target.py
  • evals/targets/eval_2_target.py
  • evals/targets/eval_3_target.py
  • evals/targets/eval_4_target.py
  • evals/targets/test/eval_1_test_target.py
  • evals/targets/test/eval_2_test_target.py
  • evals/targets/test/eval_3_test_target.py
  • evals/targets/test/eval_4_test_target.py
  • references/implementation-guide.py
  • … and 4 more

Open the folder on GitHubat commit 0e0d506

Compare with similar skills

Earth2studio Create Datasource next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Earth2studio Create Datasource compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Earth2studio Create Datasource this skillNVIDIA/skills3.5k—~2.6kAutomated safety check: PassApache-2.0
Merge Dependabot PRsonyx-dot-app/onyx32k1 repos~2.2kAutomated safety check: PassMIT
Kedro Babysitkedro-org/kedro11k—~4kAutomated safety check: PassCustom licence
Adk Setupgoogle/adk-python22k—~993Automated safety check: NotesApache-2.0
OpenROAD Issue TriageThe-OpenROAD-Project/OpenROAD3.2k—~842Automated safety check: PassBSD-3-Clause
GAIA Agent Eval Scorecardamd/gaia1.6k—~2.6kAutomated safety check: PassMIT

Similar skills

  • Merge Dependabot PRs

    onyx-dot-app/onyx

    Triages and lands a batch of open Dependabot PRs in the Onyx repo, where main is gated exclusively by GitHub's merge queue: approves and enqueues green PRs, closes superseded duplicates, fixes…

    32k GitHub starsUsed in 1 repo~2.2k tokens
    DevelopmentAuto-check passed
  • Kedro Babysit

    kedro-org/kedro

    Run Kedro's local lint / format / type-check / tests on changed files (uses the project's pre-commit hooks, ruff, mypy, pytest, lint-imports, detect-secrets, Make targets — in the right venv), or…

    11k GitHub stars~4k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Adk Setup

    google/adk-python

    Official

    Sets up a local ADK Python development environment in a git clone of the open-source adk-python repository: a uv virtual environment, all dependency extras, pre-commit hooks, and a first unit-test…

    22k GitHub stars~993 tokensUpdated today
    DevelopmentAuto-check: notes
  • OpenROAD Issue Triage

    The-OpenROAD-Project/OpenROAD

    Reproduces an OpenROAD GitHub bug from an attached tarball and shrinks the failing design with whittle.py so maintainers get a minimal test case.

    3.2k GitHub stars~842 tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Adds a release eval scorecard to a GAIA hub agent by writing a harness adapter, running a real eval, and wiring the result into the agent's README and release gate.

    1.6k GitHub stars~2.6k tokensUpdated yesterday
    DevelopmentAuto-check passed
  • Bump

    av1155/houndarr

    Bump Houndarr version and prepare a release PR. An agent skill from av1155/houndarr.

    292 GitHub stars~1.2k tokensUpdated 2 days ago
    DevelopmentAuto-check passed

More from NVIDIA/skills

All 380 skills in this repo
  • Official

    A skill your agent uses when the user wants to deploy, run, debug, tear down, or call the REST API of the RTVI-CV 2D detection / tracking microservice.

    3.5k GitHub starsUsed in 1 repo~4.5k tokens
    Auto-check passed
  • Official

    Generates, validates, compares and explains HOLOLINK_def.svh macro files for the HSB IP, using bundled Python scripts and asking before it writes anything.

    3.5k GitHub stars~2.9k tokensUpdated yesterday
    Auto-check passed
  • Official

    Runs and validates an end-to-end Mission Control demo in a locally installed Isaac Sim, with a Nova Carter robot driven through a Python server.

    3.5k GitHub stars~4.8k tokensUpdated yesterday
    Auto-check passed
  • Orchestrates defect image generation for PCBA, metal surface and glass inspection with NVIDIA Cosmos AnomalyGen on OSMO, from cold-start Day 0 to real-photo Day 1 labeling.

    3.5k GitHub stars~5k tokensUpdated yesterday
    Auto-check: notes
  • Orchestrates video data augmentation and auto-labeling workflows on OSMO, from flow selection and preflight checks to submission, monitoring and output download.

    3.5k GitHub stars~4.7k tokensUpdated yesterday
    Auto-check: notes
  • Official

    Runs NVIDIA TAO Data Services KPI analysis on object detection results, comparing predictions to ground truth and writing per-class precision, recall and AP to a CSV.

    3.5k GitHub stars~2.7k tokensUpdated yesterday
    Auto-check: notes

Works with

Questions about Earth2studio Create Datasource

What does Earth2studio Create Datasource do?

Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores. Earth2studio Create Datasource is an agent skill from NVIDIA/skills, published by the product's own GitHub organization. Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores.

When should I use Earth2studio Create Datasource?

Earth2studio Create Datasource fits situations like: fetching data with existing sources; model inference; installation tasks.

How do I install Earth2studio Create Datasource in Claude Code?

Run `npx skills add NVIDIA/skills --skill earth2studio-create-datasource -a claude-code`. Or copy the skill folder (skills/earth2studio-create-datasource in NVIDIA/skills) into .claude/skills/earth2studio-create-datasource in your project. Claude Code loads it when a task matches its description.

How do I install Earth2studio Create Datasource in Codex?

Run `npx skills add NVIDIA/skills --skill earth2studio-create-datasource -a codex`. Or copy the skill folder (skills/earth2studio-create-datasource in NVIDIA/skills) into .agents/skills/earth2studio-create-datasource in your project. Codex loads it when a task matches its description.

Can I use Earth2studio Create Datasource in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/skills --skill earth2studio-create-datasource -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/earth2studio-create-datasource, .gemini/skills/earth2studio-create-datasource, .github/skills/earth2studio-create-datasource and .opencode/skills/earth2studio-create-datasource in your project.

What does Earth2studio Create Datasource need to run?

Going by SKILL.md and its folder, Earth2studio Create Datasource needs Python and a shell for the scripts in its folder and the command-line tools its instructions call (uv, make and gh). Our summary lists: Python 3; A Bash shell.

Does Earth2studio Create Datasource access the network?

SKILL.md contains no URLs. Its commands use uv and gh, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Earth2studio Create Datasource safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Earth2studio Create Datasource use?

Earth2studio Create Datasource is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Earth2studio Create Datasource use?

About 2.6k tokens (SKILL.md is roughly 11k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 9.5k tokens, read only when the agent opens those files.

What are the alternatives to Earth2studio Create Datasource?

Skills that share tags, products or a category with Earth2studio Create Datasource: Merge Dependabot PRs (onyx-dot-app/onyx, 32k stars), Kedro Babysit (kedro-org/kedro, 11k stars), Adk Setup (google/adk-python, 22k stars) and OpenROAD Issue Triage (The-OpenROAD-Project/OpenROAD, 3.2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Earth2studio Create Datasource?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/skills, which has 3,534 GitHub stars. The repository holds 380 skills in this directory. The repository was last updated on October 7, 2026.

Source: NVIDIA/skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.