Agent skill

Geniml

by aipoch in aipoch/medical-research-skills

Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…

MITAuto-check passedResearch & Science

Install Geniml

skills CLI
$ npx skills add aipoch/medical-research-skills --skill geniml -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install aipoch/medical-research-skills geniml --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/aipoch/medical-research-skills.git skills-src && mkdir -p .claude/skills && cp -r skills-src/'scientific-skills/Data Analysis/geniml' .claude/skills/geniml && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
geniml
GitHub stars
2k
Token cost
~1.9k tokens
SKILL.md length
554 words
Files
7 (incl. references)
Skills in repo
567
Repo updated
First seen
Licence
MIT

At a glance

Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…

  • Works in 3 steps: Install → End-to-end: Build a universe → tokenize… → Single-cell ATAC-seq: tokenize cells →…
  • You need to tokenize BED collections and train embeddings for regions/cells/labels
  • SKILL.md covers When to Use, Key Features, Dependencies and Example Usage, plus 1 more section
  • Calls uv; reaches github.com

What it does

Geniml is an agent skill from aipoch/medical-research-skills. Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run similarity search and downstream ML on chromatin accessibility datasets.

Its SKILL.md is about 1.9k tokens, which your agent loads only when the skill is triggered. The skill folder holds 7 other files, including reference files (for example `geniml_audit_result_v1.json`, `references/bedspace.md` and `references/consensus_peaks.md`).

It sits in Research & Science, covering Bioinformatics, Embeddings and Vector databases. It works with Scanpy. The repository describes itself as: Hundreds of agent skills for medical research, including protocol design, data analysis, evidence insights, and academic writing. The licence is MIT.

When your agent uses it

  • You need to tokenize BED collections and train embeddings for regions/cells/labels
  • Build consensus peak universes
  • Run similarity search and downstream ML on chromatin accessibility datasets

Example prompts

  • “/geniml”

Requirements

  • Python 3

Workflow steps

3 steps, taken from the step headings in SKILL.md.

  1. Install
  2. End-to-end: Build a universe → tokenize BEDs → train Region2Vec → evaluate
  3. Single-cell ATAC-seq: tokenize cells → train scEmbed → cluster with Scanpy

What it can do on your machine

Read from SKILL.md and the folder at commit 686e09d. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Geniml loads about 1.9k tokens when it runs, and up to ~9.2k if it reads all its reference files. Until then it costs about 68 tokens; SKILL.md has 554 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~68
When it runs · the whole SKILL.md, loaded when a task matches
~1.9k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.2k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from aipoch/medical-research-skills at commit 686e09d, republished under its MIT licence (© aipoch). 554 words, ~1,851 tokens.

Download SKILL.mdSave it as .claude/skills/geniml/SKILL.md (or your agent's skills folder). This skill also uses 6 other files; get the full folder from GitHub.
name
geniml
description
Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run similarity search and downstream ML on chromatin accessibility datasets.
license
MIT
author
AIPOCH

Source: https://github.com/aipoch/medical-research-skills

When to Use

  • You have many BED files and need numeric features for clustering, similarity search, or downstream supervised learning (e.g., ChIP-seq/ATAC-seq region sets).
  • You want unsupervised embeddings of genomic regions to compare region sets across experiments (Region2Vec).
  • You need joint embeddings of regions and metadata labels (e.g., tissue/cell type/condition) to enable cross-modal queries like Region → Label or Label → Region (BEDspace).
  • You are analyzing single-cell ATAC-seq and want cell embeddings for clustering/annotation and integration with Scanpy workflows (scEmbed).
  • You need a consensus peak set (“universe”) built from multiple BED files to standardize tokenization and region definitions across datasets (Universe construction).

Key Features

  • Region2Vec: Word2vec-style unsupervised embeddings for genomic regions from tokenized BED data.
  • BEDspace: StarSpace-based joint embedding space for region sets and metadata labels; supports similarity search and cross-modal retrieval.
  • scEmbed: Single-cell ATAC-seq embedding workflow (tokenize cells → train → encode cells) compatible with Scanpy.
  • Universe (Consensus Peaks) Builder: Generates reference peak sets using multiple statistical approaches (CC, CCF, ML, HMM).
  • Utilities:
    • Tokenization: Universe-based tokenization (hard/soft tokenization patterns).
    • Evaluation: Embedding quality metrics (e.g., silhouette, Davies–Bouldin).
    • BEDshift: Region randomization/null-model generation while preserving genomic context.
    • BBClient / caching: Faster repeated access to BED resources.
    • Text2BedNN: Neural search backend for genomic queries.

Additional details are commonly documented in: references/region2vec.md, references/bedspace.md, references/scembed.md, references/consensus_peaks.md, references/utilities.md.

Dependencies

  • Python: 3.9+ (recommended)
  • geniml: latest from PyPI (or GitHub main)
  • Optional ML extras: geniml[ml] (typically pulls PyTorch and related ML dependencies)
  • Scanpy stack (for scEmbed workflows): scanpy (plus anndata, numpy, scipy)
  • StarSpace (for BEDspace training): external binary from https://github.com/facebookresearch/StarSpace
  • Universe coverage generation: uniwig (used to generate coverage tracks in universe workflows)

Example Usage

1) Install
bash
# Base install
uv pip install geniml

# With ML extras (e.g., PyTorch and related dependencies)
uv pip install "geniml[ml]"

# Development version
uv pip install git+https://github.com/databio/geniml.git
2) End-to-end: Build a universe → tokenize BEDs → train Region2Vec → evaluate
bash
# (A) Build coverage tracks (example pattern)
cat bed_files/*.bed > combined.bed
uniwig -m 25 combined.bed chrom.sizes coverage/

# (B) Build a universe (coverage cutoff method)
geniml universe build cc \
  --coverage-folder coverage/ \
  --output-file universe.bed \
  --cutoff 5 \
  --merge 100 \
  --filter-size 50
python
# (C) Tokenize BED files, train Region2Vec, and evaluate embeddings
from geniml.tokenization import hard_tokenization
from geniml.region2vec import region2vec
from geniml.evaluation import evaluate_embeddings

# 1) Tokenize BED files against the universe
hard_tokenization(
    src_folder="bed_files/",
    dst_folder="tokens/",
    universe_file="universe.bed",
    p_value_threshold=1e-9,
)

# 2) Train Region2Vec
region2vec(
    token_folder="tokens/",
    save_dir="model/",
    num_shufflings=1000,
    embedding_dim=100,
)

# 3) Evaluate (requires labels/metadata aligned to embeddings)
metrics = evaluate_embeddings(
    embeddings_file="model/embeddings.npy",
    labels_file="metadata.csv",
)

print(metrics)
3) Single-cell ATAC-seq: tokenize cells → train scEmbed → cluster with Scanpy
python
import scanpy as sc
from geniml.scembed import ScEmbed
from geniml.io import tokenize_cells

# 1) Load AnnData
adata = sc.read_h5ad("scatac_data.h5ad")

# 2) Tokenize cells using a universe
tokenize_cells(
    adata="scatac_data.h5ad",
    universe_file="universe.bed",
    output="tokens.parquet",
)

# 3) Train scEmbed
model = ScEmbed(embedding_dim=100)
model.train(dataset="tokens.parquet", epochs=100)

# 4) Encode cells and attach embeddings to AnnData
embeddings = model.encode(adata)
adata.obsm["scembed_X"] = embeddings

# 5) Standard Scanpy neighborhood graph + clustering + UMAP
sc.pp.neighbors(adata, use_rep="scembed_X")
sc.tl.leiden(adata)
sc.tl.umap(adata)

Implementation Details

Tokenization (Universe-based)
  • Goal: Convert genomic intervals into discrete “tokens” defined by a reference universe (consensus peak set).
  • Hard tokenization: Assigns intervals to universe bins/peaks deterministically (commonly used for Region2Vec/scEmbed pipelines).
  • Key parameter: p_value_threshold controls stringency of mapping/overlap significance (lower is stricter; overly strict thresholds can reduce coverage).
Show full SKILL.md (220 more words)Show less
Region2Vec (Region Embeddings)
  • Core idea: Treat each BED file (or region set) like a “document” and each universe peak like a “word”; learn embeddings using a word2vec-style objective.
  • Important knobs:
    • embedding_dim: dimensionality of learned vectors (e.g., 50–300).
    • num_shufflings: increases training signal by shuffling/co-occurrence augmentation; higher values increase runtime.
BEDspace (Joint Region + Label Embeddings)
  • Core idea: Learn a shared vector space for region sets and metadata labels using StarSpace, enabling:
    • Region → Label retrieval (predict likely labels for a query region set)
    • Label → Region retrieval (find region sets associated with a label)
  • Operational requirement: StarSpace must be installed and its path provided/configured for training.
scEmbed (Single-cell Embeddings)
  • Core idea: Apply Region2Vec-like training on tokenized single-cell accessibility profiles to produce cell embeddings.
  • Best practice: Pre-tokenize cells (e.g., to Parquet) to reduce repeated preprocessing and speed up training.
  • Downstream: Use embeddings as adata.obsm[...] and run standard Scanpy steps (neighbors, Leiden, UMAP).
Universe Construction (Consensus Peaks)
  • Purpose: Create a stable reference peak set for tokenization and cross-dataset comparability.
  • Methods:
    • CC (Coverage Cutoff): threshold-based peak calling from coverage.
    • CCF (Coverage Cutoff Flexible): cutoff with flexible boundaries/confidence intervals.
    • ML (Maximum Likelihood): probabilistic modeling of peak positions.
    • HMM (Hidden Markov Model): state-based segmentation; typically most computationally intensive.
  • Typical parameters:
    • --cutoff: minimum coverage to call peaks (CC/CCF).
    • --merge: merge distance for nearby peaks.
    • --filter-size: minimum peak length to keep.

© aipoch, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 6 other files (references) in scientific-skills/Data Analysis/geniml of aipoch/medical-research-skills.

  • SKILL.md
  • geniml_audit_result_v1.json
  • references/bedspace.md
  • references/consensus_peaks.md
  • references/region2vec.md
  • references/scembed.md
  • references/utilities.md

Open the folder on GitHubat commit 686e09d

Compare with similar skills

Geniml next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Geniml compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Geniml this skillaipoch/medical-research-skills2k—~1.9kAutomated safety check: PassMIT
Bio Data Visualization Dimensionality Reduction PlotsGPTomics/bioSkills1.2k2 repos~4.8kAutomated safety check: PassMIT
Evo2JimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
Scanpy Single-Cell Analysisdavila7/claude-code-templates32k16 repos~2.8kAutomated safety check: PassMIT
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
Genimldavila7/claude-code-templates32k12 repos~2.5kAutomated safety check: PassMIT

Similar skills

  • Produce and interpret PCA, t-SNE, UMAP, and PHATE plots for high-dimensional omics data with rigor about which method preserves what (variance, local structure, manifold, transitions)…

    1.2k GitHub starsUsed in 2 repos~4.8k tokens
    Data & AnalyticsAuto-check passed
  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Scanpy Single-Cell Analysis

    davila7/claude-code-templates

    Walks through single-cell RNA-seq analysis with Scanpy: loading .h5ad and 10X data, QC, normalization, PCA and UMAP, Leiden clustering, marker genes and cell type annotation.

    32k GitHub starsUsed in 16 repos~2.8k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Geniml

    davila7/claude-code-templates

    This skill should be used when working with genomic interval data (BED files) for machine learning tasks.

    32k GitHub starsUsed in 12 repos~2.5k tokens
    Research & ScienceAuto-check passed
  • Cellxgene Census

    davila7/claude-code-templates

    Query CZ CELLxGENE Census (61M+ cells). An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.8k tokens
    Research & ScienceAuto-check passed

More from aipoch/medical-research-skills

All 567 skills in this repo
  • Academic Poster Generator

    aipoch/medical-research-skills

    Complete workflow for generating academic research posters from PDF literature; use when you need to extract paper content from PDFs and produce a LaTeX-based poster…

    2k GitHub stars~2.2k tokensUpdated 21 days ago
    Auto-check passed
  • Diagnostic Study Quality Assessment Quadas

    aipoch/medical-research-skills

    Analyzes clinical diagnostic accuracy studies for bias using the QUADAS-2 tool.

    2k GitHub stars~1.4k tokensUpdated 21 days ago
    Auto-check passed
  • Exploratory Data Analysis

    aipoch/medical-research-skills

    Perform comprehensive exploratory data analysis on scientific data files across 200+ file formats.

    2k GitHub stars~3.7k tokensUpdated 21 days ago
    Auto-check passed
  • Iso Certification

    aipoch/medical-research-skills

    A toolkit for preparing ISO 13485:2016 certification documentation for medical device QMS.

    2k GitHub stars~1.8k tokensUpdated 21 days ago
    Auto-check passed
  • Journal Skills

    aipoch/medical-research-skills

    Recommends target journals for manuscript submission by analyzing the paper topic/abstract and the journal distribution of similar PubMed literature; use when users ask for journal…

    2k GitHub stars~1.7k tokensUpdated 21 days ago
    Auto-check passed
  • Latex Posters

    aipoch/medical-research-skills

    Creates academic-poster writing packages for LaTeX using beamerposter, tikzposter, or baposter.

    2k GitHub stars~1.3k tokensUpdated 21 days ago
    Auto-check passed

Works with

Questions about Geniml

What does Geniml do?

Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…. Geniml is an agent skill from aipoch/medical-research-skills. Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run similarity search and downstream ML on chromatin accessibility datasets.

When should I use Geniml?

Geniml fits situations like: you need to tokenize BED collections and train embeddings for regions/cells/labels; build consensus peak universes; run similarity search and downstream ML on chromatin accessibility datasets.

How do I install Geniml in Claude Code?

Run `npx skills add aipoch/medical-research-skills --skill geniml -a claude-code`. Or copy the skill folder (scientific-skills/Data Analysis/geniml in aipoch/medical-research-skills) into .claude/skills/geniml in your project. Claude Code loads it when a task matches its description.

How do I install Geniml in Codex?

Run `npx skills add aipoch/medical-research-skills --skill geniml -a codex`. Or copy the skill folder (scientific-skills/Data Analysis/geniml in aipoch/medical-research-skills) into .agents/skills/geniml in your project. Codex loads it when a task matches its description.

Can I use Geniml in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add aipoch/medical-research-skills --skill geniml -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/geniml, .gemini/skills/geniml, .github/skills/geniml and .opencode/skills/geniml in your project.

What does Geniml need to run?

Going by SKILL.md and its folder, Geniml needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Geniml access the network?

SKILL.md names 1 domain. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. This is read from the text; nothing was executed.

Is Geniml safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Geniml use?

Geniml is published under the MIT licence (declared in SKILL.md). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Geniml use?

About 1.9k tokens (SKILL.md is roughly 7.4k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 7.3k tokens, read only when the agent opens those files.

What are the alternatives to Geniml?

Skills that share tags, products or a category with Geniml: Bio Data Visualization Dimensionality Reduction Plots (GPTomics/bioSkills, 1.2k stars), Evo2 (JimLiu/science-skills, 227 stars), Scanpy Single-Cell Analysis (davila7/claude-code-templates, 32k stars) and Scgpt (JimLiu/science-skills, 227 stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Geniml?

aipoch (a GitHub organization) maintains it in aipoch/medical-research-skills, which has 1,974 GitHub stars. The repository holds 567 skills in this directory. The repository was last updated on September 17, 2026.

Source: aipoch/medical-research-skills on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.