Install the "lamindb-data-management" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/lamindb-data-management into .claude/skills/lamindb-data-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lamindb-data-management", then confirm the skill loads.
Claude Code copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Type this inside Codex. $skill-installer <name> installs a curated skill from openai/skills. The installer writes to $CODEX_HOME/skills (default ~/.codex/skills). Restart Codex if the skill does not show up.
skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a codex
Project install goes to .agents/skills/; add -g for ~/.codex/skills/.
Install the "lamindb-data-management" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/lamindb-data-management into .agents/skills/lamindb-data-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lamindb-data-management", then confirm the skill loads.
Codex copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a cursor
Project install goes to .agents/skills/; add -g for ~/.cursor/skills/.
Install the "lamindb-data-management" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/lamindb-data-management into .cursor/skills/lamindb-data-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lamindb-data-management", then confirm the skill loads.
Cursor copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
--scope user (default) or --scope workspace; --path is the subfolder of the repo that holds the skill; --consent skips the security confirmation prompt.
skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a gemini-cli
Project install goes to .agents/skills/; add -g for ~/.gemini/skills/.
Install the "lamindb-data-management" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/lamindb-data-management into .gemini/skills/lamindb-data-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lamindb-data-management", then confirm the skill loads.
Gemini CLI copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Installs for Copilot at project scope by default; add --scope user for a personal install. Preview a skill first with gh skill preview. Needs GitHub CLI 2.90.0 or later (public preview).
skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a github-copilot
Project install goes to .agents/skills/; add -g for ~/.copilot/skills/.
Install the "lamindb-data-management" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/lamindb-data-management into .github/skills/lamindb-data-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lamindb-data-management", then confirm the skill loads.
GitHub Copilot copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
skills CLI
$ npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a opencode
OpenCode documents no install command of its own. Project install goes to .agents/skills/; add -g for ~/.config/opencode/skills/.
Install the "lamindb-data-management" agent skill from https://github.com/jaechang-hits/SciAgent-Skills/tree/main/skills/systems-biology-multiomics/lamindb-data-management into .opencode/skills/lamindb-data-management/ in this project. Copy the whole folder (SKILL.md and every file beside it), keep the folder name "lamindb-data-management", then confirm the skill loads.
OpenCode copies the folder itself, the same result as the manual copy. Check what it changed before you commit it.
Facts
Skill name
lamindb-data-management
GitHub stars
371
Used in
2 other repos
Token cost
~4k tokens
SKILL.md length
773 words
Files
1
Skills in repo
169
Repo updated
First seen
Licence
Apache-2.0
At a glance
Open-source FAIR biology data framework. An agent skill from jaechang-hits/SciAgent-Skills.
Works in 6 steps: Artifacts — Data Objects → Lineage Tracking → Querying and Filtering → …
Tasks that involve Bioinformatics
SKILL.md covers Overview, When to Use, Prerequisites and Quick Start, plus 9 more sections
Calls pip
What it does
Lamindb Data Management is an agent skill from jaechang-hits/SciAgent-Skills. Open-source FAIR biology data framework. Version artifacts (AnnData, DataFrame, Zarr), track lineage, validate via ontologies (Bionty), query datasets. Integrates with Nextflow, Snakemake, W&B, scVI. For scRNA-seq use scanpy; for ontology lookups use bionty.
Its SKILL.md is about 4k tokens, which your agent loads only when the skill is triggered. It is a single SKILL.md file with no bundled scripts.
It sits in Research & Science, covering Bioinformatics, Reproducible research and Knowledge graphs. It works with Nextflow, AnnData, Scanpy and Zarr. The repository describes itself as: 197 bioinformatics & life science skills for Claude Code and AI agents — BixBench 92.0% accuracy. RNA-seq, single-cell, drug discovery, proteomics, and more. Powers OmicsHorizon. The licence is Apache-2.0.
When your agent uses it
Tasks that involve Bioinformatics
Tasks that involve Reproducible research
Tasks that involve Knowledge graphs
Example prompts
“/lamindb-data-management”
Requirements
Python 3
Workflow steps
6 steps, taken from the step headings in SKILL.md.
Read from SKILL.md and the folder at commit 82c862c. It shows what the files ask for, not the result of running them.
Tool permissions
Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.
From allowed-tools in the SKILL.md frontmatter.
Runs code
Shell commands in SKILL.md call:
pip
From the folder's file list and the shell code blocks in SKILL.md.
Network
Links to these hosts (documentation or services it may open):
docs.lamin.ai
github.com
From URLs in SKILL.md, links to its own repository left out.
Credentials
Names no API keys, tokens, secrets or passwords.
From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.
Context cost
Lamindb Data Management loads about 4k tokens when it runs. Until then it costs about 71 tokens; SKILL.md has 773 words of instructions outside code blocks.
Always· name and description, kept in context so the agent knows when to use it
~71
When it runs· the whole SKILL.md, loaded when a task matches
~4k
Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.
Safety
Auto-check passed
The automated check found no risky patterns in SKILL.md.
Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.
Download SKILL.mdSave it as .claude/skills/lamindb-data-management/SKILL.md (or your agent's skills folder).
name
lamindb-data-management
description
Open-source FAIR biology data framework. Version artifacts (AnnData, DataFrame, Zarr), track lineage, validate via ontologies (Bionty), query datasets. Integrates with Nextflow, Snakemake, W&B, scVI. For scRNA-seq use scanpy; for ontology lookups use bionty.
license
Apache-2.0
LaminDB — Biological Data Management
Overview
LaminDB is an open-source data framework for biology that makes data queryable, traceable, and FAIR (Findable, Accessible, Interoperable, Reusable). It combines data lakehouse architecture, lineage tracking, biological ontology validation, and a unified Python API for managing biological datasets from raw files to annotated, curated artifacts.
When to Use
Managing and versioning biological datasets (scRNA-seq, spatial, flow cytometry, multi-modal)
Tracking computational lineage (which code produced which data)
Validating and curating data against biological ontologies (cell types, genes, tissues, diseases)
Building queryable data lakehouses across multiple experiments
Ensuring reproducibility with automatic environment and provenance capture
Integrating with workflow managers (Nextflow, Snakemake) or MLOps (W&B, MLflow)
Standardizing metadata with ontology-based annotation (Bionty)
For single-cell analysis pipelines (clustering, DE), use scanpy instead
For ontology lookups only without data management, use bionty directly
Prerequisites
bash
pip install lamindb
# With extras for specific data types
pip install 'lamindb[bionty,zarr,fcs]'
Setup: Requires instance initialization before use:
Always wrap analysis with ln.track() / ln.finish(): This captures lineage automatically. Without it, artifacts have no provenance.
Use hierarchical keys: Structure as project/experiment/datatype/file.ext (e.g., immunology/exp42/scrna/counts.h5ad). This enables prefix-based queries.
Anti-pattern — duplicating data instead of versioning: Use the revises= parameter to create new versions, not new keys for the same dataset.
Validate early: Run schema validation before analysis. Catching bad metadata early saves debugging time downstream.
Use ontologies for standardization: Map free-text labels to ontology terms (e.g., "T helper cell" → CL:0000912). This enables cross-dataset queries.
Anti-pattern — loading large files without checking size: Use .filter().df() to inspect metadata first, then .load() or .open() (backed mode) for large files.
Query metadata first, load data second: Filter with .filter() to find relevant artifacts, then load only what you need.
Common Recipes
Recipe: Bulk Dataset Registration
python
import lamindb as ln
from pathlib import Path
ln.track()
data_dir = Path("raw_data/")
for fcs_file in data_dir.glob("*.fcs"):
artifact = ln.Artifact(str(fcs_file), key=f"flow_cytometry/{fcs_file.name}").save()
artifact.features.add_values({"assay": "flow_cytometry", "source": "batch_import"})
print(f"Registered: {fcs_file.name} -> {artifact.uid}")
ln.finish()
Recipe: View and Export Lineage
python
import lamindb as ln
artifact = ln.Artifact.get(key="results/final_analysis.h5ad")
# View lineage graph (opens in browser or notebook)
artifact.view_lineage()
# Programmatic lineage access
run = artifact.run
print(f"Created by: {run.transform.name}")
print(f"User: {run.created_by.name}")
print(f"Date: {run.created_at}")
print(f"Input artifacts: {[a.key for a in run.input_artifacts.all()]}")
Recipe: Ontology Hierarchy Exploration
python
import bionty as bt
bt.CellType.import_source()
t_cell = bt.CellType.get(name="T cell")
# Explore hierarchy
print(f"Parents: {[p.name for p in t_cell.parents.all()]}")
print(f"Children: {[c.name for c in t_cell.children.all()]}")
# Find all descendants
descendants = t_cell.children.all()
for child in descendants:
grandchildren = child.children.all()
print(f" {child.name}: {[gc.name for gc in grandchildren]}")
Troubleshooting
Problem
Cause
Solution
InstanceNotSetupError
Instance not initialized
Run lamin init --storage ./data --name my-project
ln.track() fails
No transform context
Run inside a notebook/script, not REPL; or pass transform explicitly
Artifact key conflict
Key already exists (not a version)
Use revises= for versioning, or choose a different key
ValidationError
Data doesn't match schema
Run curator.validate() to see specific failures; standardize terms
Slow queries on large instances
No index on filtered field
Use .df() for overview first; add database indexes for frequently filtered fields
Ontology import fails
Network issue or wrong organism
Check internet connection; specify organism="human" explicitly
FileNotFoundError on .cache()
Cloud artifact not synced
Check storage connectivity; use artifact.load() instead for in-memory access
Related Skills
anndata-data-structure — AnnData format used as primary data container in LaminDB for single-cell data
scanpy-scrna-seq — single-cell analysis pipeline; LaminDB manages data that scanpy analyzes
scvi-tools-single-cell — deep learning models for single-cell; integrates with LaminDB for data/model tracking
We found 2 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 2 other GitHub owners. This page covers the copy in jaechang-hits/SciAgent-Skills, which our catalogue first saw on October 7, 2026.
Lamindb Data Management next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.
Lamindb Data Management compared with similar skills
Skill
Stars
Used in
Tokens
Auto-check
Licence
Repo updated
Lamindb Data Management this skilljaechang-hits/SciAgent-Skills
Stores and operates on sparse expression matrices for single-cell and large bulk RNA-seq, covering dgCMatrix/dgRMatrix/dgTMatrix when-each-is-fast, the dgCMatrix (CSC, R) <- CSR (Python) implicit…
Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.
Manages biological datasets with LaminDB: versioned artifacts, run lineage, ontology-based annotation, schema validation and links to workflow managers and ML tools.
This skill should be used when working with annotated data matrices in Python, particularly for single-cell genomics analysis, managing experimental measurements with metadata, or handling…
Open-source FAIR biology data framework. An agent skill from jaechang-hits/SciAgent-Skills. Lamindb Data Management is an agent skill from jaechang-hits/SciAgent-Skills. Open-source FAIR biology data framework.
When should I use Lamindb Data Management?
Lamindb Data Management fits situations like: tasks that involve Bioinformatics; tasks that involve Reproducible research; tasks that involve Knowledge graphs.
How do I install Lamindb Data Management in Claude Code?
Run `npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a claude-code`. Or copy the skill folder (skills/systems-biology-multiomics/lamindb-data-management in jaechang-hits/SciAgent-Skills) into .claude/skills/lamindb-data-management in your project. Claude Code loads it when a task matches its description.
How do I install Lamindb Data Management in Codex?
Run `npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a codex`. Or copy the skill folder (skills/systems-biology-multiomics/lamindb-data-management in jaechang-hits/SciAgent-Skills) into .agents/skills/lamindb-data-management in your project. Codex loads it when a task matches its description.
Can I use Lamindb Data Management in Cursor, Gemini CLI or GitHub Copilot?
Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jaechang-hits/SciAgent-Skills --skill lamindb-data-management -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/lamindb-data-management, .gemini/skills/lamindb-data-management, .github/skills/lamindb-data-management and .opencode/skills/lamindb-data-management in your project.
What does Lamindb Data Management need to run?
Going by SKILL.md and its folder, Lamindb Data Management needs the command-line tools its instructions call (pip). Our summary lists: Python 3.
Does Lamindb Data Management access the network?
SKILL.md names 2 domains. As links in the text: docs.lamin.ai and github.com. This is read from the text; nothing was executed.
Is Lamindb Data Management safe to install?
Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.
What licence does Lamindb Data Management use?
Lamindb Data Management is published under the Apache-2.0 licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.
How many tokens does Lamindb Data Management use?
About 4k tokens (SKILL.md is roughly 16k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full.
What are the alternatives to Lamindb Data Management?
Skills that share tags, products or a category with Lamindb Data Management: Bio Expression Matrix Sparse Handling (GPTomics/bioSkills, 1.2k stars), Anndata (K-Dense-AI/scientific-agent-skills, 48k stars), Lamindb (aipoch/medical-research-skills, 2k stars) and Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.
Who maintains Lamindb Data Management?
jaechang-hits (a GitHub user) maintains it in jaechang-hits/SciAgent-Skills, which has 371 GitHub stars. The repository holds 169 skills in this directory. The repository was last updated on September 29, 2026.
Source: jaechang-hits/SciAgent-Skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.