Agent skill

Optimuskg

by mims-harvard in mims-harvard/OptimusKG

Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client.

MITAuto-check passedData & Analytics

Install Optimuskg

skills CLI
$ npx skills add mims-harvard/OptimusKG --skill optimuskg -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install mims-harvard/OptimusKG optimuskg --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/mims-harvard/OptimusKG.git skills-src && mkdir -p .claude/skills && cp -r skills-src/skills/optimuskg .claude/skills/optimuskg && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
optimuskg
GitHub stars
147
Token cost
~1.9k tokens
SKILL.md length
581 words
Files
2
Skills in repo
4
Repo updated
First seen
Licence
MIT

At a glance

Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client.

  • Tasks that involve DataFrames
  • SKILL.md covers When to use, Installation, Quick start and Choosing the right loader, plus 6 more sections
  • Calls uv and pip; reaches dataverse.harvard.edu
  • Tasks that involve Knowledge graphs

What it does

Optimuskg is an agent skill from mims-harvard/OptimusKG. Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client. Use this when loading, querying, filtering, or analyzing OptimusKG data — genes, drugs, diseases, phenotypes, anatomy, pathways, and their relationships — as Polars DataFrames or a NetworkX graph, or when downloading the published graph from Harvard Dataverse.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 2 other files (for example `reference/graph-schema.md`).

It sits in Data & Analytics, covering DataFrames and Knowledge graphs. It works with NetworkX, Polars and Python. The repository describes itself as: A modern multimodal knowledge graph with type-specific metadata across biomedical domains. The licence is MIT.

When your agent uses it

  • Tasks that involve DataFrames
  • Tasks that involve Knowledge graphs

Example prompts

  • “/optimuskg”

Requirements

  • Python 3

What it can do on your machine

Read from SKILL.md and the folder at commit 4fb3529. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv
    • pip

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • dataverse.harvard.edu

    Also links to:

    • optimuskg.ai
    • doi.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Optimuskg loads about 1.9k tokens when it runs. Until then it costs about 92 tokens; SKILL.md has 581 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~92
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from mims-harvard/OptimusKG at commit 4fb3529, republished under its MIT licence (© mims-harvard). 581 words, ~1,872 tokens.

Download SKILL.mdSave it as .claude/skills/optimuskg/SKILL.md (or your agent's skills folder). This skill also uses 1 other file; get the full folder from GitHub.
name
optimuskg
description
Guide for using OptimusKG, the biomedical knowledge graph, through the `optimuskg` Python client. Use this when loading, querying, filtering, or analyzing OptimusKG data — genes, drugs, diseases, phenotypes, anatomy, pathways, and their relationships — as Polars DataFrames or a NetworkX graph, or when downloading the published graph from Harvard Dataverse.

OptimusKG

OptimusKG is a modern multimodal biomedical knowledge graph (190,531 nodes across 10 entity types, 21,813,816 edges across 27 relation types) integrating 65 resources grounded in 18 ontologies via the BioCypher framework and Biolink Model. It is published on Harvard Dataverse as Apache Parquet files and consumed through the optimuskg PyPI client.

This skill covers using the published graph via the client. It is not for developing the data pipeline in the mims-harvard/optimuskg repo — that work uses the repo's own /node-catalog-sync skill.

When to use

Use the optimuskg client when you need to:

  • Download OptimusKG node/edge tables for analysis.
  • Load the graph as Polars DataFrames or a NetworkX MultiDiGraph.
  • Filter by entity type (Gene, Drug, Disease, …) or relation type.
  • Build biomedical analyses, Graph-RAG, or ML over the graph.

Always prefer this client over hand-rolling Dataverse downloads — it resolves file IDs automatically and caches locally.

Installation

bash
uv add optimuskg     # in a uv project (preferred — see the `uv` skill)
pip install optimuskg

The client depends on polars (DataFrames) and networkx (graph view).

Quick start

python
import optimuskg

# Download a specific file; returns its local cached path
local_path = optimuskg.get_file("nodes/gene.parquet")

# Read a single Parquet file as a Polars DataFrame
drugs = optimuskg.load_parquet("nodes/drug.parquet")

# Load nodes + edges as Polars DataFrames (lcc=True -> largest connected component only)
nodes, edges = optimuskg.load_graph(lcc=True)

# Load as a NetworkX MultiDiGraph with properties merged onto node/edge attrs
G = optimuskg.load_networkx(lcc=True)

Choosing the right loader

NeedFunctionReturns
Just the file on diskget_file(path, *, force=False)pathlib.Path
One table as a DataFrameload_parquet(path, *, force=False, **read_parquet_kwargs)pl.DataFrame
Whole graph as DataFramesload_graph(*, lcc=False, force=False)(nodes, edges) tuple of pl.DataFrame
Whole graph as NetworkXload_networkx(*, lcc=False, force=False, parse_properties=True)nx.MultiDiGraph

Notes:

  • load_parquet forwards extra kwargs to pl.read_parquet — e.g. push down column selection: optimuskg.load_parquet("nodes/drug.parquet", columns=["id"]).
  • force=True re-downloads even if the file is already cached.
  • load_networkx always builds a MultiDiGraph regardless of edge directionality; call G.to_undirected() if you need an undirected view.
  • Memory: loading the full graph into NetworkX needs several GB (190k nodes, 21M edges) and emits a warning. Pass lcc=True for a smaller, connected variant unless you specifically need every node.

File paths

Paths mirror the catalog layout under data/gold/kg/parquet/ in the source repo:

python
optimuskg.get_file("nodes.parquet")                              # full nodes table
optimuskg.get_file("edges.parquet")                              # full edges table
optimuskg.get_file("largest_connected_component_nodes.parquet")  # LCC nodes
optimuskg.get_file("largest_connected_component_edges.parquet")  # LCC edges
optimuskg.get_file("nodes/gene.parquet")                         # only Gene (GEN) nodes
optimuskg.get_file("edges/disease_gene.parquet")                 # only DIS-GEN edges

nodes/<type>.parquet files use the lowercase entity name (gene, drug, …); edges/<a>_<b>.parquet files use the lowercase node pair (disease_gene, drug_gene, …).

Show full SKILL.md (277 more words)Show less

Graph schema (at a glance)

Node table columns: id, label (the type code, e.g. GEN), properties. Edge table columns: from, to, label (e.g. DIS-GEN), relation (e.g. ASSOCIATED_WITH), undirected, properties.

In the unified nodes.parquet / edges.parquet tables, properties is a JSON string. In the stratified per-type files (nodes/<type>.parquet, edges/<label>.parquet) it is expanded into native typed columns as a Polars Struct.

The 10 node type codes:

LabelTypeLabelType
GENGeneANAAnatomy
DISDiseaseMFNMolecular Function
BPOBiological ProcessCCOCellular Component
PHEPhenotypePWYPathway
DRGDrugEXPExposure

For the full edge-label/relation taxonomy (all 27 edge types and their relation strings) and per-type property fields, see reference/graph-schema.md.

Common patterns

Filter Polars DataFrames by type/relation:

python
import polars as pl

nodes, edges = optimuskg.load_graph(lcc=True)
genes = nodes.filter(pl.col("label") == "GEN")
dis_gen = edges.filter(pl.col("relation") == "ASSOCIATED_WITH")

Filter a NetworkX graph (properties are merged onto attrs):

python
G = optimuskg.load_networkx(lcc=True)

# Nodes by type code
genes = [n for n, a in G.nodes(data=True) if a["label"] == "GEN"]

# Edges by relation
expression = [
    (u, v) for u, v, a in G.edges(data=True)
    if a["relation"] == "EXPRESSION_PRESENT"
]

Pass parse_properties=False to load_networkx to keep properties as a raw JSON string instead of merging parsed keys into the attribute dicts.

Configuration

The client targets doi:10.7910/DVN/IYNGEV on https://dataverse.harvard.edu by default, and caches downloads in platformdirs.user_cache_dir("optimuskg") (~/.cache/optimuskg on Linux, ~/Library/Caches/optimuskg on macOS). Cache keys include the dataset version, so a new release invalidates it automatically.

Override from code or via environment variables:

python
optimuskg.set_cache_dir("/data/optimuskg-cache")
optimuskg.set_doi("doi:10.7910/DVN/EXAMPLE")     # target a different release
optimuskg.set_server("https://dataverse.example.org")  # non-Harvard installation

# Read current settings
optimuskg.get_cache_dir(); optimuskg.get_doi(); optimuskg.get_server()
bash
export OPTIMUSKG_CACHE_DIR=/data/optimuskg-cache
export OPTIMUSKG_DOI=doi:10.7910/DVN/EXAMPLE
export OPTIMUSKG_SERVER=https://dataverse.example.org

Citing and license

  • Cite OptimusKG when you use it in research — see https://optimuskg.ai/docs/citation and the dataset DOI 10.7910/DVN/IYNGEV.
  • License: the OptimusKG codebase is MIT. The integrated source datasets keep their own licenses and terms of use, which may restrict redistribution or commercial use of a given graph subset — review each source's terms. See https://optimuskg.ai/docs/license.

Documentation

For full details (function signatures, per-type schemas, edge relations), read the official docs — they expose machine-readable text files:

© mims-harvard, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 1 other file in skills/optimuskg of mims-harvard/OptimusKG.

  • SKILL.md
  • reference/graph-schema.md

Open the folder on GitHubat commit 4fb3529

Compare with similar skills

Optimuskg next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Optimuskg compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Optimuskg this skillmims-harvard/OptimusKG147—~1.9kAutomated safety check: PassMIT
Hybrid-Engine Data Analysiscode-yeongyu/oh-my-openagent70k—~1.4kAutomated safety check: PassCustom licence
Narwhalsanam-org/metaxy124—~3.3kAutomated safety check: PassApache-2.0
Optimize For GPUK-Dense-AI/scientific-agent-skills48k1 repos~3.4kAutomated safety check: PassMIT
PolarsK-Dense-AI/scientific-agent-skills48k1 repos~3.3kAutomated safety check: PassMIT
Ingesting Dataancoleman/ai-design-components525—~1.9kAutomated safety check: PassMIT

Similar skills

  • Hybrid-Engine Data Analysis

    code-yeongyu/oh-my-openagent

    Analyzes CSV, Parquet and JSON data with DuckDB, Polars, numpy and matplotlib, preferring a persistent kernel over repeated one-shot processes.

    70k GitHub stars~1.4k tokensUpdated today
    Data & AnalyticsAuto-check passed
  • Narwhals

    anam-org/metaxy

    Effectively use Narwhals to write dataframe-agnostic code that works seamlessly across multiple Python dataframe libraries.

    124 GitHub stars~3.3k tokensUpdated 10 days ago
    Data & AnalyticsAuto-check passed
  • Optimize For GPU

    K-Dense-AI/scientific-agent-skills

    GPU-accelerates scientific Python on NVIDIA hardware and verifies that the result is correct and faster.

    48k GitHub starsUsed in 1 repo~3.4k tokens
    Data & AnalyticsAuto-check passed
  • Polars

    K-Dense-AI/scientific-agent-skills

    High-performance DataFrame library for Python ETL, analytics, and pandas migration.

    48k GitHub starsUsed in 1 repo~3.3k tokens
    Data & AnalyticsAuto-check passed
  • Ingesting Data

    ancoleman/ai-design-components

    Data ingestion patterns for loading data from cloud storage, APIs, files, and streaming sources into databases.

    525 GitHub stars~1.9k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed
  • Transforming Data

    ancoleman/ai-design-components

    Transform raw data into analytical assets using ETL/ELT patterns, SQL (dbt), Python (pandas/polars/PySpark), and orchestration (Airflow).

    525 GitHub stars~3k tokensUpdated 10 mo ago
    Data & AnalyticsAuto-check passed

More from mims-harvard/OptimusKG

  • Scientific Visualization

    mims-harvard/OptimusKG

    Create publication figures with matplotlib/seaborn/plotly. An agent skill from mims-harvard/OptimusKG.

    147 GitHub starsUsed in 19 repos~6.3k tokens
    Auto-check passed
  • Impeccable

    mims-harvard/OptimusKG

    Create distinctive, production-grade frontend interfaces with high design quality.

    147 GitHub starsUsed in 3 repos~5.5k tokens
    Auto-check passed
  • Node Catalog Sync

    mims-harvard/OptimusKG

    Enforce synchronization between Kedro node files and catalog YAML files in the OptimusKG project.

    147 GitHub stars~823 tokensUpdated 19 days ago
    Auto-check passed

Questions about Optimuskg

What does Optimuskg do?

Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client. Optimuskg is an agent skill from mims-harvard/OptimusKG. Guide for using OptimusKG, the biomedical knowledge graph, through the optimuskg Python client.

When should I use Optimuskg?

Optimuskg fits situations like: tasks that involve DataFrames; tasks that involve Knowledge graphs.

How do I install Optimuskg in Claude Code?

Run `npx skills add mims-harvard/OptimusKG --skill optimuskg -a claude-code`. Or copy the skill folder (skills/optimuskg in mims-harvard/OptimusKG) into .claude/skills/optimuskg in your project. Claude Code loads it when a task matches its description.

How do I install Optimuskg in Codex?

Run `npx skills add mims-harvard/OptimusKG --skill optimuskg -a codex`. Or copy the skill folder (skills/optimuskg in mims-harvard/OptimusKG) into .agents/skills/optimuskg in your project. Codex loads it when a task matches its description.

Can I use Optimuskg in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add mims-harvard/OptimusKG --skill optimuskg -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/optimuskg, .gemini/skills/optimuskg, .github/skills/optimuskg and .opencode/skills/optimuskg in your project.

What does Optimuskg need to run?

Going by SKILL.md and its folder, Optimuskg needs the command-line tools its instructions call (uv and pip). Our summary lists: Python 3.

Does Optimuskg access the network?

SKILL.md names 3 domains. In commands or code: dataverse.harvard.edu; the agent is likely to contact it when it follows the instructions. As links in the text: optimuskg.ai and doi.org. This is read from the text; nothing was executed.

Is Optimuskg safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Optimuskg use?

Optimuskg is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Optimuskg use?

About 1.9k tokens (SKILL.md is roughly 7.5k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.

What are the alternatives to Optimuskg?

Skills that share tags, products or a category with Optimuskg: Hybrid-Engine Data Analysis (code-yeongyu/oh-my-openagent, 70k stars), Narwhals (anam-org/metaxy, 124 stars), Optimize For GPU (K-Dense-AI/scientific-agent-skills, 48k stars) and Polars (K-Dense-AI/scientific-agent-skills, 48k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Optimuskg?

mims-harvard (a GitHub organization) maintains it in mims-harvard/OptimusKG, which has 147 GitHub stars. The repository holds 4 skills in this directory. The repository was last updated on September 21, 2026.

Source: mims-harvard/OptimusKG on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.