Official agent skill

Physicsnemo Cfd Create Dataset Adapter

by NVIDIA in NVIDIA/physicsnemo-cfd

Create a new dataset adapter for the PhysicsNeMo CFD benchmarking workflow.

OfficialApache-2.0Auto-check passedResearch & Science

Install Physicsnemo Cfd Create Dataset Adapter

skills CLI
$ npx skills add NVIDIA/physicsnemo-cfd --skill physicsnemo-cfd-create-dataset-adapter -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install NVIDIA/physicsnemo-cfd physicsnemo-cfd-create-dataset-adapter --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/NVIDIA/physicsnemo-cfd.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/physicsnemo-cfd-create-dataset-adapter .claude/skills/physicsnemo-cfd-create-dataset-adapter && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
physicsnemo-cfd-create-dataset-adapter
GitHub stars
153
Token cost
~3k tokens
SKILL.md length
956 words
Files
5
Skills in repo
3
Repo updated
First seen
Licence
Apache-2.0

At a glance

Create a new dataset adapter for the PhysicsNeMo CFD benchmarking workflow.

  • Works in 5 steps: Explore the new dataset → Write the adapter class → Register and test → …
  • The user wants to add a new CFD dataset
  • SKILL.md covers Reference files to read first, Step 1: Explore the new dataset, Step 2: Write the adapter class and Step 3: Register and test, plus 4 more sections
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md

What it does

Physicsnemo Cfd Create Dataset Adapter is an agent skill from NVIDIA/physicsnemo-cfd, published by the product's own GitHub organization. Create a new dataset adapter for the PhysicsNeMo CFD benchmarking workflow. Use when the user wants to add a new CFD dataset, write a DatasetAdapter, integrate a new mesh format, or benchmark models on custom data.

Its SKILL.md is about 3k tokens, which your agent loads only when the skill is triggered. The skill folder holds 5 other files (for example `BENCHMARK.md`, `evals/evals.json` and `skill-card.md`).

It sits in Research & Science, covering Physical and earth sciences. It works with NVIDIA AI Platform. The repository describes itself as: Library for using the models trained in PhysicsNeMo in Engineering and CFD workflows. The licence is Apache-2.0.

When your agent uses it

  • The user wants to add a new CFD dataset
  • Write a DatasetAdapter
  • Integrate a new mesh format
  • Benchmark models on custom data

Example prompts

  • “/physicsnemo-cfd-create-dataset-adapter”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Explore the new dataset
  2. Write the adapter class
  3. Register and test
  4. Run inference and benchmark
  5. Make permanent (optional)

What it can do on your machine

Read from SKILL.md and the folder at commit 0612ec4. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md (its code samples are python).

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Physicsnemo Cfd Create Dataset Adapter loads about 3k tokens when it runs. Until then it costs about 63 tokens; SKILL.md has 956 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~63
When it runs · the whole SKILL.md, loaded when a task matches
~3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from NVIDIA/physicsnemo-cfd at commit 0612ec4, republished under its Apache-2.0 licence (© NVIDIA). 956 words, ~3,017 tokens.

Download SKILL.mdSave it as .claude/skills/physicsnemo-cfd-create-dataset-adapter/SKILL.md (or your agent's skills folder). This skill also uses 4 other files; get the full folder from GitHub.
name
physicsnemo-cfd-create-dataset-adapter
description
Create a new dataset adapter for the PhysicsNeMo CFD benchmarking workflow. Use when the user wants to add a new CFD dataset, write a DatasetAdapter, integrate a new mesh format, or benchmark models on custom data.
license
Apache-2.0

Create a Dataset Adapter

Guide the user through adding a new CFD dataset to the benchmarking workflow by writing a DatasetAdapter subclass.

Reference files to read first

Before starting, read these files for context:

  • physicsnemo/cfd/evaluation/datasets/adapter_registry.py — base class and registry
  • physicsnemo/cfd/evaluation/datasets/schema.py — CanonicalCase and build_predictions_dict
  • physicsnemo/cfd/evaluation/datasets/adapters/drivaerml.py — reference adapter implementation
  • workflows/benchmarking/notebooks/adding_a_new_dataset.ipynb — end-to-end tutorial (writes a DrivAerStar adapter: format conversion, field renaming, WSS sign flip, STL creation)

Step 1: Explore the new dataset

Ask the user for the dataset path, then inspect one file. Report not just array names but their component count, dtype, and value range, plus mesh.bounds and any geometry arrays — the decision table below needs all of these:

python
import numpy as np
import pyvista as pv

mesh = pv.read("<path_to_one_file>")
print(f"Type: {type(mesh).__name__}, Points: {mesh.n_points}, Cells: {mesh.n_cells}")
print(f"Bounds (xmin,xmax,ymin,ymax,zmin,zmax): {mesh.bounds}")
for loc, data in [("cell", mesh.cell_data), ("point", mesh.point_data)]:
    for name in data.keys():
        arr = np.asarray(data[name])
        comps = arr.shape[1] if arr.ndim > 1 else 1
        print(f"  [{loc}] {name}: comps={comps}, dtype={arr.dtype}, "
              f"range=({arr.min():.3g}, {arr.max():.3g})")
# Explicit geometry arrays some datasets ship (DrivAerML has none):
print("Has Normals:", "Normals" in mesh.cell_data or "Normals" in mesh.point_data)
print("Has Area:", "Area" in mesh.cell_data or "Area" in mesh.point_data)

Identify these differences from the canonical schema:

QuestionWhat to look for
File format.vtp, .vtu, .vtk, or a non-VTK format (CGNS, OpenFOAM, HDF5, CSV, ...)? Model wrappers ultimately read .vtp (surface) or .vtu (volume) XML — see "Reading non-PyVista source formats".
Directory layoutFlat directory? Nested run_<id>/ dirs? How are case IDs derived from filenames?
Pressure field nameThe canonical key is pressure. What is the VTK array name?
WSS field nameThe canonical key is shear_stress (N, 3). Is it a single vector or separate scalar components?
Sign conventionsCompare field ranges with DrivAerML. Are normals, WSS, or pressure flipped?
Extra arraysAre there explicit Normals or Area arrays? DrivAerML has none — remove them if present.
STL filesAre separate STL geometry files available? If not, the surface mesh itself is the geometry.
Coordinate frame & scaleCompare mesh.bounds and units against the training dataset. Matters only for geometry-referenced checkpoints (e.g. DrivAerML-trained). See "Match geometry orientation and scale".
Inference domainSurface (.vtp) or volume (.vtu)?

Step 2: Write the adapter class

Subclass DatasetAdapter with these methods:

python
from pathlib import Path
from physicsnemo.cfd.evaluation.datasets.adapter_registry import DatasetAdapter, register_adapter
from physicsnemo.cfd.evaluation.datasets.schema import CanonicalCase

class MyDatasetAdapter(DatasetAdapter):
    def __init__(self, root: str, **kwargs):
        self._root = Path(root)

    @classmethod
    def inference_domain_from_kwargs(cls, kwargs=None):
        return "surface"  # or "volume"

    def list_cases(self):
        # Return list of case ID strings
        ...

    def load_case(self, case_id: str) -> CanonicalCase:
        # 1. Read the mesh file
        # 2. Build ground_truth dict with canonical keys:
        #    - "pressure": np.float32 array
        #    - "shear_stress": np.float32 array of shape (N, 3)
        #    For volume: "pressure", "velocity" (N,3), "turbulent_viscosity"
        # 3. Return CanonicalCase(case_id, mesh_path, mesh_type, ground_truth, inference_domain)
        ...
Map source arrays to canonical keys

ground_truth must use the framework's canonical keys, but source files rarely use those names. The canonical vocabulary (see schema.py / build_predictions_dict) is:

Canonical keyShapeDomain
pressure(N,)surface, volume
shear_stress(N, 3)surface
velocity(N, 3)volume
turbulent_viscosity(N,)volume

Build an explicit rename map from the source names you found in Step 1:

python
RENAME = {"pMean": "pressure", "wallShearStress": "shear_stress"}
ground_truth = {
    canon: np.asarray(mesh.cell_data[src], dtype=np.float32)
    for src, canon in RENAME.items()
}

When names are ambiguous, disambiguate by: component count (a 3-comp field is velocity or shear_stress), dtype/value range, and — decisively — what the model's training data called each field (see "Why conventions must match the training data"). Do not confuse this source→canonical map with the separate canonical→VTK-name map in output.mesh_field_names (Step 4), which controls the written arrays.

Common transformations in load_case

Reading non-PyVista source formats: pv.read handles VTK-family files, but CFD ground truth often ships as CGNS, OpenFOAM cases, Ensight, Tecplot, HDF5/.npz, or CSV point clouds. Only reading changes — the target is still a canonical .vtp/.vtu mesh plus a ground_truth dict:

python
# meshio covers many formats (CGNS, Ensight, ...); wrap to PyVista:
import meshio, pyvista as pv
mesh = pv.wrap(meshio.read(src_path))

# OpenFOAM case directory:
mesh = pv.OpenFOAMReader(case_foam_file).read()

# Raw arrays (HDF5 / npz / CSV): build the mesh, then attach fields:
cloud = pv.PolyData(points_xyz)          # (N, 3) float array
cloud["pressure"] = p_values             # attach source arrays

Format conversion (legacy .vtk → .vtp):

python
mesh = pv.read(vtk_path).extract_surface()
mesh.save(vtp_path)

Combining separate WSS scalars into a vector:

python
wss = np.stack([mesh.cell_data["WSSx"], mesh.cell_data["WSSy"], mesh.cell_data["WSSz"]], axis=1)

Removing explicit Normals/Area (DrivAerML convention):

python
for key in ["Normals", "Area"]:
    if key in mesh.cell_data:
        del mesh.cell_data[key]

Creating STL from surface mesh (when no STL is shipped):

python
mesh.extract_surface().triangulate().save(stl_path)

The STL must be named drivaer_{int(case_id)}.stl in the same directory as the VTP for the model wrappers to find it.

Show full SKILL.md (468 more words)Show less
Match geometry orientation and scale

Geometry-referenced models (e.g. DoMINO) normalize the mesh/STL coordinates against a fixed bounding box baked into the checkpoint from its training dataset: DoMINO reads cfg.data.bounding_box_surface.min/max (and bounding_box.min/max for volume) and maps every coordinate into that box. If the new dataset's geometry sits in a different frame, origin, or unit scale, it lands in the wrong normalized space — predictions are wrong even when field names and signs are correct.

This only matters when the checkpoint was trained on a specific geometry-referenced dataset (e.g. DrivAerML). For scale/translation-invariant models, or when the model was trained on this same dataset, skip it.

Match three things to the training dataset (DrivAerML reference bounds below, in meters, from the DoMINO config):

Boxmin (x, y, z)max (x, y, z)
Surface-1.5, -1.4, -0.325.0, 1.4, 1.4
Volume-3.5, -2.25, -0.328.5, 2.25, 3.00
  • Orientation / axes: same convention — x streamwise (length), y width, z up. Permute or rotate if the new data uses a different up-axis or flipped sign.
  • Origin / position: the bounding box should start near the same (x, y, z) minimum, so the geometry falls inside the model's domain box.
  • Scale / units: extents must be the same order of magnitude. Millimetre data must be scaled to meters (×0.001).

Check mesh.bounds and transform in load_case before saving the prepared VTP/STL:

python
b = mesh.bounds  # (xmin, xmax, ymin, ymax, zmin, zmax)
# ~1000x larger extents => mm; scale to meters. A swapped axis range => reorient.
mesh.points *= 0.001
mesh.points += np.array([x_off, y_off, z_off], dtype=np.float32)  # translate to match origin
Caching pattern

Do expensive conversions lazily and cache:

python
def _prepare_case(self, case_id):
    prepared_path = self._root / "_prepared" / f"{case_id}.vtp"
    if not prepared_path.exists():
        # ... convert and save
    return str(prepared_path)

Step 3: Register and test

python
register_adapter("my_dataset", MyDatasetAdapter)

adapter = MyDatasetAdapter(root="/path/to/data")
cases = adapter.list_cases()
case = adapter.load_case(cases[0])
assert case.ground_truth is not None
assert "pressure" in case.ground_truth

Step 4: Run inference and benchmark

Build a config and run:

python
from physicsnemo.cfd.evaluation.config import Config
from physicsnemo.cfd.evaluation.benchmarks.engine import run_benchmark

config = Config.from_dict({
    "run": {"device": "cuda:0", "output_dir": "results"},
    "model": {"name": "<model_name>", "inference_domain": "<surface|volume>", ...},
    "dataset": {"name": "my_dataset", "root": "/path/to/data", "case_ids": cases[:2]},
    "output": {
        "ground_truth_mesh_field_names": {"pressure": "<vtk_gt_name>", "shear_stress": "<vtk_gt_name>"},
        "mesh_field_names": {"pressure": "<vtk_pred_name>", "shear_stress": "<vtk_pred_name>"},
    },
    "metrics": ["l2_pressure", "l2_shear_stress", "drag", "lift"],
    "reports": {"enabled": False},
})
results = run_benchmark(config)

Step 5: Make permanent (optional)

Save the adapter to physicsnemo/cfd/evaluation/datasets/adapters/<name>.py and register in adapters/__init__.py:

python
from physicsnemo.cfd.evaluation.datasets.adapters.<name> import MyDatasetAdapter
register_adapter("my_dataset", MyDatasetAdapter)

Why conventions must match the training data

The field name mappings, sign conventions, and format conversions in the adapter exist because the model checkpoint was trained on a specific dataset (e.g., DrivAerML) with specific conventions. The adapter bridges the gap between the new dataset's conventions and the training data's conventions — not some abstract standard. If a model is retrained directly on the new dataset, the adapter would not need these transformations. When writing an adapter, always ask: "What conventions did the model's training data use?" and map to those.

Gotchas

  • DistributedManager: Model wrappers call DistributedManager.initialize(). In notebooks without torchrun, set env vars first: WORLD_SIZE=1, RANK=0, LOCAL_RANK=0, MASTER_ADDR=localhost, MASTER_PORT=12355.
  • STL naming: DoMINO looks for drivaer_{tag}.stl, GeoTransolver looks for drivaer_{tag}_single_solid.stl then *.stl. Both now fall back to any *.stl in the directory.
  • VTP vs VTK: Model wrappers use VTK XML readers internally. Legacy .vtk files must be converted to .vtp/.vtu.
  • Checkpoint loading: Some wrappers need trusted_torch_load_context() for PyTorch 2.6+ checkpoint compatibility.
  • Domain-scoped metrics: l2_pressure resolves to different implementations for surface vs volume based on inference_domain. Use the same metric name for both.
  • Geometry frame: geometry-referenced checkpoints (DoMINO) assume the training dataset's coordinate frame and scale. A mm-vs-m or flipped-axis mismatch produces wrong predictions with no error raised. See "Match geometry orientation and scale".

© NVIDIA, Apache-2.0. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 4 other files in skills/physicsnemo-cfd-create-dataset-adapter of NVIDIA/physicsnemo-cfd.

  • SKILL.md
  • BENCHMARK.md
  • evals/evals.json
  • skill-card.md
  • skill.oms.sig

Open the folder on GitHubat commit 0612ec4

Compare with similar skills

Physicsnemo Cfd Create Dataset Adapter next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Physicsnemo Cfd Create Dataset Adapter compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Physicsnemo Cfd Create Dataset Adapter this skillNVIDIA/physicsnemo-cfd153—~3kAutomated safety check: PassApache-2.0
AstropyzLanqing/codex-claude-academic-skills4.6k14 repos~2.9kAutomated safety check: PassBSD-3-Clause
PymatgenzLanqing/codex-claude-academic-skills4.6k12 repos~5kAutomated safety check: PassMIT
Cantera Ignition DelayK-Dense-AI/scientific-agent-skills48k1 repos~2.2kAutomated safety check: PassMIT
Weathertrpc-group/trpc-agent-go1.8k8 repos~591Automated safety check: PassApache-2.0
Pymol VisualizationChatMol/ChatMol372—~1.2kAutomated safety check: PassMIT

Similar skills

  • Astropy

    zLanqing/codex-claude-academic-skills

    Comprehensive Python library for astronomy and astrophysics.

    4.6k GitHub starsUsed in 14 repos~2.9k tokens
    Research & ScienceAuto-check passed
  • Pymatgen

    zLanqing/codex-claude-academic-skills

    Materials science toolkit. An agent skill from zLanqing/codex-claude-academic-skills.

    4.6k GitHub starsUsed in 12 repos~5k tokens
    Research & ScienceAuto-check passed
  • Cantera Ignition Delay

    K-Dense-AI/scientific-agent-skills

    Runs Cantera constant-volume or constant-pressure ignition simulations and reports temperature-based ignition delay with mechanism provenance and checks.

    48k GitHub starsUsed in 1 repo~2.2k tokens
    Research & ScienceAuto-check passed
  • Weather

    trpc-group/trpc-agent-go

    Get current weather and forecasts via wttr.in or Open-Meteo.

    1.8k GitHub starsUsed in 8 repos~591 tokens
    Research & ScienceAuto-check passed
  • Pymol Visualization

    ChatMol/ChatMol

    Generate publication-quality molecular visualization images using PyMOL.

    372 GitHub stars~1.2k tokensUpdated 6 mo ago
    Research & ScienceAuto-check passed
  • Complexa Binder Design

    NVIDIA-BioNeMo/bionemo-agent-toolkit

    Run a complete protein binder design campaign with NVIDIA Proteina-Complexa: resolve a target structure and hotspots from a name/sequence/PDB, co-design binder sequence+structure with reward-guided…

    478 GitHub stars~3.1k tokensUpdated today
    Research & ScienceAuto-check: notes

More from NVIDIA/physicsnemo-cfd

  • Official

    Create a new model wrapper for the PhysicsNeMo CFD benchmarking workflow.

    153 GitHub stars~5.1k tokensUpdated 1 mo ago
    Auto-check passed
  • Official

    Create a custom metric for the PhysicsNeMo CFD benchmarking workflow.

    153 GitHub stars~1.6k tokensUpdated 1 mo ago
    Auto-check passed

Questions about Physicsnemo Cfd Create Dataset Adapter

What does Physicsnemo Cfd Create Dataset Adapter do?

Create a new dataset adapter for the PhysicsNeMo CFD benchmarking workflow. Physicsnemo Cfd Create Dataset Adapter is an agent skill from NVIDIA/physicsnemo-cfd, published by the product's own GitHub organization. Create a new dataset adapter for the PhysicsNeMo CFD benchmarking workflow.

When should I use Physicsnemo Cfd Create Dataset Adapter?

Physicsnemo Cfd Create Dataset Adapter fits situations like: the user wants to add a new CFD dataset; write a DatasetAdapter; integrate a new mesh format; benchmark models on custom data.

How do I install Physicsnemo Cfd Create Dataset Adapter in Claude Code?

Run `npx skills add NVIDIA/physicsnemo-cfd --skill physicsnemo-cfd-create-dataset-adapter -a claude-code`. Or copy the skill folder (skills/physicsnemo-cfd-create-dataset-adapter in NVIDIA/physicsnemo-cfd) into .claude/skills/physicsnemo-cfd-create-dataset-adapter in your project. Claude Code loads it when a task matches its description.

How do I install Physicsnemo Cfd Create Dataset Adapter in Codex?

Run `npx skills add NVIDIA/physicsnemo-cfd --skill physicsnemo-cfd-create-dataset-adapter -a codex`. Or copy the skill folder (skills/physicsnemo-cfd-create-dataset-adapter in NVIDIA/physicsnemo-cfd) into .agents/skills/physicsnemo-cfd-create-dataset-adapter in your project. Codex loads it when a task matches its description.

Can I use Physicsnemo Cfd Create Dataset Adapter in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add NVIDIA/physicsnemo-cfd --skill physicsnemo-cfd-create-dataset-adapter -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/physicsnemo-cfd-create-dataset-adapter, .gemini/skills/physicsnemo-cfd-create-dataset-adapter, .github/skills/physicsnemo-cfd-create-dataset-adapter and .opencode/skills/physicsnemo-cfd-create-dataset-adapter in your project.

What does Physicsnemo Cfd Create Dataset Adapter need to run?

SKILL.md names no scripts, command-line tools or credentials: Physicsnemo Cfd Create Dataset Adapter is instructions for the agent only. Our summary lists: Python 3.

Does Physicsnemo Cfd Create Dataset Adapter access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is Physicsnemo Cfd Create Dataset Adapter safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Physicsnemo Cfd Create Dataset Adapter use?

Physicsnemo Cfd Create Dataset Adapter is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Physicsnemo Cfd Create Dataset Adapter use?

About 3k tokens (SKILL.md is roughly 12k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Physicsnemo Cfd Create Dataset Adapter?

Skills that share tags, products or a category with Physicsnemo Cfd Create Dataset Adapter: Astropy (zLanqing/codex-claude-academic-skills, 4.6k stars), Pymatgen (zLanqing/codex-claude-academic-skills, 4.6k stars), Cantera Ignition Delay (K-Dense-AI/scientific-agent-skills, 48k stars) and Weather (trpc-group/trpc-agent-go, 1.8k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Physicsnemo Cfd Create Dataset Adapter?

NVIDIA (a GitHub organization, an official publisher) maintains it in NVIDIA/physicsnemo-cfd, which has 153 GitHub stars. The repository holds 3 skills in this directory. The repository was last updated on August 18, 2026.

Source: NVIDIA/physicsnemo-cfd on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.