Agent skill

Node Catalog Sync

by mims-harvard in mims-harvard/OptimusKG

Enforce synchronization between Kedro node files and catalog YAML files in the OptimusKG project.

MITAuto-check passedData & Analytics

Install Node Catalog Sync

skills CLI
$ npx skills add mims-harvard/OptimusKG --skill node-catalog-sync -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mims-harvard/OptimusKG node-catalog-sync --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mims-harvard/OptimusKG.git skills-src && mkdir -p .claude/skills && cp -r skills-src/.agents/skills/node-catalog-sync .claude/skills/node-catalog-sync && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
node-catalog-sync
GitHub stars
146
Token cost
~823 tokens
SKILL.md length
288 words
Files
1
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Enforce synchronization between Kedro node files and catalog YAML files in the OptimusKG project.

  • Works in 4 steps: Identify affected catalog entries → Update dataset ID and filepath → Rerun node and sync catalog → …
  • Editing any Python node file under optimuskg/pipelines//nodes/
  • SKILL.md covers Path Mapping, Sync Workflow and Catalog YAML Structure
  • Calls uv

What it does

Node Catalog Sync is an agent skill from mims-harvard/OptimusKG. Enforce synchronization between Kedro node files and catalog YAML files in the OptimusKG project. Use when editing any Python node file under optimuskg/pipelines//nodes/, including modifying run() functions, changing node() inputs/outputs, adding/removing/renaming DataFrame columns, or changing column types. Also use when creating new nodes or deleting existing ones.

Its SKILL.md is about 820 tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.

It sits in Data & Analytics, covering DataFrames. It works with Python. The repository describes itself as: A modern multimodal knowledge graph with type-specific metadata across biomedical domains. The licence is MIT.

When your agent uses it

  • Editing any Python node file under optimuskg/pipelines//nodes/
  • Including modifying run() functions
  • Changing node() inputs/outputs
  • Adding/removing/renaming DataFrame columns

Example prompts

  • “/node-catalog-sync”

Requirements

  • Python 3

Workflow steps

4 steps, taken from the step headings in SKILL.md.

  1. Identify affected catalog entries
  2. Update dataset ID and filepath
  3. Rerun node and sync catalog
  4. Cascade downstream

What it can do on your machine

Read from SKILL.md and the folder at commit 4fb3529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md. Its commands use uv, which can reach the network depending on how they are called.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Node Catalog Sync loads about 823 tokens when it runs. Until then it costs about 97 tokens; SKILL.md has 288 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~97
When it runs · the whole SKILL.md, loaded when a task matches
~823

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mims-harvard/OptimusKG at commit 4fb3529, republished under its MIT licence (© mims-harvard). 288 words, ~823 tokens.

Download SKILL.mdSave it as .claude/skills/node-catalog-sync/SKILL.md (or your agent's skills folder).
name
node-catalog-sync
description
Enforce synchronization between Kedro node files and catalog YAML files in the OptimusKG project. Use when editing any Python node file under optimuskg/pipelines/*/nodes/, including modifying run() functions, changing node() inputs/outputs, adding/removing/renaming DataFrame columns, or changing column types. Also use when creating new nodes or deleting existing ones.

Node–Catalog Sync

When a node file is edited, the corresponding catalog YAML files must be updated to stay in sync.

Path Mapping

Node outputs are namespace-prefixed by pipeline.py. Use this table to locate catalog files:

LayerNode outputsCatalog IDYAML pathfilepath
Bronze"{src}.{name}"bronze.{src}.{name}conf/base/catalog/bronze/{src}/{name}.ymldata/bronze/{src}/{name}.parquet
Silver nodes"nodes.{entity}"silver.nodes.{entity}conf/base/catalog/silver/nodes/{entity}.ymldata/silver/nodes/{entity}.parquet
Silver edges"edges.{e1}_{e2}"silver.edges.{e1}_{e2}conf/base/catalog/silver/edges/{e1}_{e2}.ymldata/silver/edges/{e1}_{e2}.parquet
Gold"kg.{fmt}"gold.kg.{fmt}conf/base/catalog/gold/{fmt}.ymldata/gold/kg/{fmt}/

Multiple outputs (list) produce one YAML file per output.

Sync Workflow

Follow all 4 steps in order after every node file edit.

Step 1: Identify affected catalog entries
  1. Read the node() call in the edited file to find outputs (string or list).
  2. Determine the namespace from the corresponding pipeline.py (e.g., namespace="bronze").
  3. Construct the full catalog ID: {namespace}.{output}.
  4. Locate the YAML file using the path mapping table above.
Step 2: Update dataset ID and filepath

Only if the outputs value in the node() call changed:

  1. Update the YAML top-level key to the new full catalog ID.
  2. Update filepath following the convention in the path mapping table.
  3. Rename the YAML file to match the new output name.
Step 3: Rerun node and sync catalog
  1. Rerun the node: uv run kedro run --nodes={node_name}
  2. Sync schema and checksum: uv run cli sync-catalog --dataset {catalog_id}
    • This reads the parquet file on disk and updates both load_args.schema and metadata.checksum in the YAML automatically.
    • Use --dry-run to preview changes first.
    • Use --validate to check without writing (useful in CI).
  3. Never delete the checksum property.
Step 4: Cascade downstream
  1. List all downstream nodes using DryRunner (no execution, just shows the DAG):
    uv run kedro run --from-nodes={node_name} --runner=optimuskg.runners.DryRunner
  2. Rerun the edited node and all its downstream dependents:
    uv run kedro run --from-nodes={node_name}
  3. Sync all affected catalog entries:
    uv run cli sync-catalog

Catalog YAML Structure

yaml
{catalog_id}:
  type: optimuskg.datasets.polars.ParquetDataset
  filepath: data/{layer}/{path}.parquet
  load_args:
    schema:
      column_name: pl.Type
      struct_column:
        nested_field: pl.Type
  metadata:
    checksum: {blake2b_hex_digest}
    kedro-viz:
      layer: {layer}

© mims-harvard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

Just SKILL.md in .agents/skills/node-catalog-sync of mims-harvard/OptimusKG.

Open the folder on GitHubat commit 4fb3529

Compare with similar skills

Node Catalog Sync next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Node Catalog Sync compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Node Catalog Sync this skillmims-harvard/OptimusKG146—~823Automated safety check: PassMIT
Chdb Datastorevemetric/vemetric3942 repos~1.4kAutomated safety check: PassApache-2.0
Polar Python SDKpolarsource/polar10k—~1.8kAutomated safety check: PassApache-2.0
CSV Data Summarizercoffeefuelbump/csv-data-summarizer-claude-skill4682 repos~1.4kAutomated safety check: PassNone
Python Executorcortega26/chile-hub1132 repos~1.5kAutomated safety check: PassMIT
Retentioneering Product Analyticsretentioneering/retentioneering-tools920—~1.6kAutomated safety check: PassApache-2.0

Similar skills

  • Chdb Datastore

    vemetric/vemetric

    A skill your agent uses when the user has tabular data (pandas DataFrame, parquet, csv, Arrow, json) and wants to filter, group, aggregate, join, or speed up slow pandas.

    394 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Polar Python SDK

    polarsource/polar

    Integrate Polar billing in server-side Python applications using the versioned Polar and PolarAsync clients.

    10k GitHub stars~1.8k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • CSV Data Summarizer

    coffeefuelbump/csv-data-summarizer-claude-skill

    Analyzes CSV files, generates summary stats, and plots quick visualizations using Python and pandas.

    468 GitHub starsUsed in 2 repos~1.4k tokens
    Data & AnalyticsAuto-check passed
  • Python Executor

    cortega26/chile-hub

    Execute Python code in a safe sandboxed environment via [inference.sh](https://inference.sh).

    113 GitHub starsUsed in 2 repos~1.5k tokens
    Data & AnalyticsAuto-check passed
  • Retentioneering Product Analytics

    retentioneering/retentioneering-tools

    Analyze event logs, clickstreams, user paths, product funnels, retention, behavioral segments, transition graphs, step matrices, sequence patterns, and customer journeys using Retentioneering.

    920 GitHub stars~1.6k tokensUpdated yesterday
    Data & AnalyticsAuto-check passed
  • Pandas Pro

    Jeffallan/claude-skills

    Handles pandas DataFrame work: cleaning, merging, groupby aggregation, pivots, time-series resampling and memory tuning, with checks on dtypes, shapes and nulls.

    12k GitHub starsUsed in 1 repo~1.5k tokens
    Data & AnalyticsAuto-check passed

More from mims-harvard/OptimusKG

  • Scientific Visualization

    mims-harvard/OptimusKG

    Create publication figures with matplotlib/seaborn/plotly. An agent skill from mims-harvard/OptimusKG.

    146 GitHub starsUsed in 19 repos~6.3k tokens
    Auto-check passed
  • Impeccable

    mims-harvard/OptimusKG

    Create distinctive, production-grade frontend interfaces with high design quality.

    146 GitHub starsUsed in 3 repos~5.5k tokens
    Auto-check passed
  • Optimuskg

    mims-harvard/OptimusKG

    Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client.

    146 GitHub stars~1.9k tokensUpdated 16 days ago
    Auto-check passed

Works with

Questions about Node Catalog Sync

What does Node Catalog Sync do?

Enforce synchronization between Kedro node files and catalog YAML files in the OptimusKG project. Node Catalog Sync is an agent skill from mims-harvard/OptimusKG. Enforce synchronization between Kedro node files and catalog YAML files in the OptimusKG project.

When should I use Node Catalog Sync?

Node Catalog Sync fits situations like: editing any Python node file under optimuskg/pipelines//nodes/; including modifying run() functions; changing node() inputs/outputs; adding/removing/renaming DataFrame columns.

How do I install Node Catalog Sync in Claude Code?

Run `npx skills add mims-harvard/OptimusKG --skill node-catalog-sync -a claude-code`. Or copy the skill folder (.agents/skills/node-catalog-sync in mims-harvard/OptimusKG) into .claude/skills/node-catalog-sync in your project. Claude Code loads it when a task matches its description.

How do I install Node Catalog Sync in Codex?

Run `npx skills add mims-harvard/OptimusKG --skill node-catalog-sync -a codex`. Or copy the skill folder (.agents/skills/node-catalog-sync in mims-harvard/OptimusKG) into .agents/skills/node-catalog-sync in your project. Codex loads it when a task matches its description.

Can I use Node Catalog Sync in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mims-harvard/OptimusKG --skill node-catalog-sync -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/node-catalog-sync, .gemini/skills/node-catalog-sync, .github/skills/node-catalog-sync and .opencode/skills/node-catalog-sync in your project.

What does Node Catalog Sync need to run?

Going by SKILL.md and its folder, Node Catalog Sync needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Node Catalog Sync access the network?

SKILL.md contains no URLs. Its commands use uv, which can reach the network depending on how they are called. This is read from the text; nothing was executed.

Is Node Catalog Sync safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Node Catalog Sync use?

Node Catalog Sync is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Node Catalog Sync use?

About 823 tokens (SKILL.md is roughly 3.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Node Catalog Sync?

Skills that share tags, products or a category with Node Catalog Sync: Chdb Datastore (vemetric/vemetric, 394 stars), Polar Python SDK (polarsource/polar, 10k stars), CSV Data Summarizer (coffeefuelbump/csv-data-summarizer-claude-skill, 468 stars) and Python Executor (cortega26/chile-hub, 113 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Node Catalog Sync?

mims-harvard (a GitHub organization) maintains it in mims-harvard/OptimusKG, which has 146 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on September 21, 2026.

Source: mims-harvard/OptimusKG on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.