Agent skill

Geniml

by davila7 in davila7/claude-code-templates

This skill should be used when working with genomic interval data (BED files) for machine learning tasks.

MITAuto-check passedResearch & Science

Install Geniml

skills CLI
$ npx skills add davila7/claude-code-templates --skill geniml -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install davila7/claude-code-templates geniml --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/davila7/claude-code-templates.git skills-src && mkdir -p .claude/skills && cp -r skills-src/cli-tool/components/skills/scientific/geniml .claude/skills/geniml && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
geniml
GitHub stars
32k
Used in
11 other repos
Token cost
~2.5k tokens
SKILL.md length
823 words
Files
6 (incl. references)
Skills in repo
478
Repo updated
First seen
Licence
MIT

At a glance

This skill should be used when working with genomic interval data (BED files) for machine learning tasks.

  • Works in 5 steps: Region2Vec: Genomic Region Embeddings → BEDspace: Joint Region and Metadata… → scEmbed: Single-Cell Chromatin… → …
  • Training region embeddings (Region2Vec
  • SKILL.md covers Overview, Installation, Core Capabilities and Common Workflows, plus 6 more sections
  • Calls uv; reaches github.com

What it does

Geniml is an agent skill from davila7/claude-code-templates. This skill should be used when working with genomic interval data (BED files) for machine learning tasks. Use for training region embeddings (Region2Vec, BEDspace), single-cell ATAC-seq analysis (scEmbed), building consensus peaks (universes), or any ML-based analysis of genomic regions. Applies to BED file collections, scATAC-seq data, chromatin accessibility datasets, and region-based genomic feature learning.

Its SKILL.md is about 2.5k tokens, which your agent loads only when the skill is triggered. The skill folder holds 6 other files, including reference files (for example `references/bedspace.md`, `references/consensus_peaks.md` and `references/region2vec.md`).

It sits in Research & Science, covering Bioinformatics and Embeddings. The repository describes itself as: CLI tool for configuring and monitoring Claude Code. The licence is MIT.

When your agent uses it

  • Training region embeddings (Region2Vec
  • Single-cell ATAC-seq analysis (scEmbed)
  • Building consensus peaks (universes)
  • Any ML-based analysis of genomic regions

Example prompts

  • “/geniml”

Requirements

  • Python 3

Workflow steps

5 steps, taken from the step headings in SKILL.md.

  1. Region2Vec: Genomic Region Embeddings
  2. BEDspace: Joint Region and Metadata Embeddings
  3. scEmbed: Single-Cell Chromatin Accessibility Embeddings
  4. Consensus Peaks: Universe Building
  5. Utilities: Supporting Tools

What it can do on your machine

Read from SKILL.md and the folder at commit 46b4d8b. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    Shell commands in SKILL.md call:

    • uv

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    Hosts in commands or code, which the agent is likely to contact:

    • github.com

    Also links to:

    • docs.bedbase.org

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

Geniml loads about 2.5k tokens when it runs, and up to ~9.3k if it reads all its reference files. Until then it costs about 106 tokens; SKILL.md has 823 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~106
When it runs · the whole SKILL.md, loaded when a task matches
~2.5k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~9.3k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from davila7/claude-code-templates at commit 46b4d8b, republished under its MIT licence (© davila7). 823 words, ~2,504 tokens.

Download SKILL.mdSave it as .claude/skills/geniml/SKILL.md (or your agent's skills folder). This skill also uses 5 other files; get the full folder from GitHub.
name
geniml
description
This skill should be used when working with genomic interval data (BED files) for machine learning tasks. Use for training region embeddings (Region2Vec, BEDspace), single-cell ATAC-seq analysis (scEmbed), building consensus peaks (universes), or any ML-based analysis of genomic regions. Applies to BED file collections, scATAC-seq data, chromatin accessibility datasets, and region-based genomic feature learning.

Geniml: Genomic Interval Machine Learning

Overview

Geniml is a Python package for building machine learning models on genomic interval data from BED files. It provides unsupervised methods for learning embeddings of genomic regions, single cells, and metadata labels, enabling similarity searches, clustering, and downstream ML tasks.

Installation

Install geniml using uv:

bash
uv uv pip install geniml

For ML dependencies (PyTorch, etc.):

bash
uv uv pip install 'geniml[ml]'

Development version from GitHub:

bash
uv uv pip install git+https://github.com/databio/geniml.git

Core Capabilities

Geniml provides five primary capabilities, each detailed in dedicated reference files:

1. Region2Vec: Genomic Region Embeddings

Train unsupervised embeddings of genomic regions using word2vec-style learning.

Use for: Dimensionality reduction of BED files, region similarity analysis, feature vectors for downstream ML.

Workflow:

  1. Tokenize BED files using a universe reference
  2. Train Region2Vec model on tokens
  3. Generate embeddings for regions

Reference: See references/region2vec.md for detailed workflow, parameters, and examples.

2. BEDspace: Joint Region and Metadata Embeddings

Train shared embeddings for region sets and metadata labels using StarSpace.

Use for: Metadata-aware searches, cross-modal queries (region→label or label→region), joint analysis of genomic content and experimental conditions.

Workflow:

  1. Preprocess regions and metadata
  2. Train BEDspace model
  3. Compute distances
  4. Query across regions and labels

Reference: See references/bedspace.md for detailed workflow, search types, and examples.

3. scEmbed: Single-Cell Chromatin Accessibility Embeddings

Train Region2Vec models on single-cell ATAC-seq data for cell-level embeddings.

Use for: scATAC-seq clustering, cell-type annotation, dimensionality reduction of single cells, integration with scanpy workflows.

Workflow:

  1. Prepare AnnData with peak coordinates
  2. Pre-tokenize cells
  3. Train scEmbed model
  4. Generate cell embeddings
  5. Cluster and visualize with scanpy

Reference: See references/scembed.md for detailed workflow, parameters, and examples.

4. Consensus Peaks: Universe Building

Build reference peak sets (universes) from BED file collections using multiple statistical methods.

Use for: Creating tokenization references, standardizing regions across datasets, defining consensus features with statistical rigor.

Workflow:

  1. Combine BED files
  2. Generate coverage tracks
  3. Build universe using CC, CCF, ML, or HMM method

Methods:

  • CC (Coverage Cutoff): Simple threshold-based
  • CCF (Coverage Cutoff Flexible): Confidence intervals for boundaries
  • ML (Maximum Likelihood): Probabilistic modeling of positions
  • HMM (Hidden Markov Model): Complex state modeling

Reference: See references/consensus_peaks.md for method comparison, parameters, and examples.

5. Utilities: Supporting Tools

Additional tools for caching, randomization, evaluation, and search.

Available utilities:

  • BBClient: BED file caching for repeated access
  • BEDshift: Randomization preserving genomic context
  • Evaluation: Metrics for embedding quality (silhouette, Davies-Bouldin, etc.)
  • Tokenization: Region tokenization utilities (hard, soft, universe-based)
  • Text2BedNN: Neural search backends for genomic queries

Reference: See references/utilities.md for detailed usage of each utility.

Common Workflows

Basic Region Embedding Pipeline
python
from geniml.tokenization import hard_tokenization
from geniml.region2vec import region2vec
from geniml.evaluation import evaluate_embeddings

# Step 1: Tokenize BED files
hard_tokenization(
    src_folder='bed_files/',
    dst_folder='tokens/',
    universe_file='universe.bed',
    p_value_threshold=1e-9
)

# Step 2: Train Region2Vec
region2vec(
    token_folder='tokens/',
    save_dir='model/',
    num_shufflings=1000,
    embedding_dim=100
)

# Step 3: Evaluate
metrics = evaluate_embeddings(
    embeddings_file='model/embeddings.npy',
    labels_file='metadata.csv'
)
scATAC-seq Analysis Pipeline
python
import scanpy as sc
from geniml.scembed import ScEmbed
from geniml.io import tokenize_cells

# Step 1: Load data
adata = sc.read_h5ad('scatac_data.h5ad')

# Step 2: Tokenize cells
tokenize_cells(
    adata='scatac_data.h5ad',
    universe_file='universe.bed',
    output='tokens.parquet'
)

# Step 3: Train scEmbed
model = ScEmbed(embedding_dim=100)
model.train(dataset='tokens.parquet', epochs=100)

# Step 4: Generate embeddings
embeddings = model.encode(adata)
adata.obsm['scembed_X'] = embeddings

# Step 5: Cluster with scanpy
sc.pp.neighbors(adata, use_rep='scembed_X')
sc.tl.leiden(adata)
sc.tl.umap(adata)
Universe Building and Evaluation
bash
# Generate coverage
cat bed_files/*.bed > combined.bed
uniwig -m 25 combined.bed chrom.sizes coverage/

# Build universe with coverage cutoff
geniml universe build cc \
  --coverage-folder coverage/ \
  --output-file universe.bed \
  --cutoff 5 \
  --merge 100 \
  --filter-size 50

# Evaluate universe quality
geniml universe evaluate \
  --universe universe.bed \
  --coverage-folder coverage/ \
  --bed-folder bed_files/

CLI Reference

Geniml provides command-line interfaces for major operations:

bash
# Region2Vec training
geniml region2vec --token-folder tokens/ --save-dir model/ --num-shuffle 1000

# BEDspace preprocessing
geniml bedspace preprocess --input regions/ --metadata labels.csv --universe universe.bed

# BEDspace training
geniml bedspace train --input preprocessed.txt --output model/ --dim 100

# BEDspace search
geniml bedspace search -t r2l -d distances.pkl -q query.bed -n 10

# Universe building
geniml universe build cc --coverage-folder coverage/ --output universe.bed --cutoff 5

# BEDshift randomization
geniml bedshift --input peaks.bed --genome hg38 --preserve-chrom --iterations 100
Show full SKILL.md (396 more words)Show less

When to Use Which Tool

Use Region2Vec when:

  • Working with bulk genomic data (ChIP-seq, ATAC-seq, etc.)
  • Need unsupervised embeddings without metadata
  • Comparing region sets across experiments
  • Building features for downstream supervised learning

Use BEDspace when:

  • Metadata labels available (cell types, tissues, conditions)
  • Need to query regions by metadata or vice versa
  • Want joint embedding space for regions and labels
  • Building searchable genomic databases

Use scEmbed when:

  • Analyzing single-cell ATAC-seq data
  • Clustering cells by chromatin accessibility
  • Annotating cell types from scATAC-seq
  • Integration with scanpy is desired

Use Universe Building when:

  • Need reference peak sets for tokenization
  • Combining multiple experiments into consensus
  • Want statistically rigorous region definitions
  • Building standard references for a project

Use Utilities when:

  • Need to cache remote BED files (BBClient)
  • Generating null models for statistics (BEDshift)
  • Evaluating embedding quality (Evaluation)
  • Building search interfaces (Text2BedNN)

Best Practices

General Guidelines
  • Universe quality is critical: Invest time in building comprehensive, well-constructed universes
  • Tokenization validation: Check coverage (>80% ideal) before training
  • Parameter tuning: Experiment with embedding dimensions, learning rates, and training epochs
  • Evaluation: Always validate embeddings with multiple metrics and visualizations
  • Documentation: Record parameters and random seeds for reproducibility
Performance Considerations
  • Pre-tokenization: For scEmbed, always pre-tokenize cells for faster training
  • Memory management: Large datasets may require batch processing or downsampling
  • Computational resources: ML/HMM universe methods are computationally intensive
  • Model caching: Use BBClient to avoid repeated downloads
Integration Patterns
  • With scanpy: scEmbed embeddings integrate seamlessly as adata.obsm entries
  • With BEDbase: Use BBClient for accessing remote BED repositories
  • With Hugging Face: Export trained models for sharing and reproducibility
  • With R: Use reticulate for R integration (see utilities reference)

Geniml is part of the BEDbase ecosystem:

  • BEDbase: Unified platform for genomic regions
  • BEDboss: Processing pipeline for BED files
  • Gtars: Genomic tools and utilities
  • BBClient: Client for BEDbase repositories

Additional Resources

Troubleshooting

"Tokenization coverage too low":

  • Check universe quality and completeness
  • Adjust p-value threshold (try 1e-6 instead of 1e-9)
  • Ensure universe matches genome assembly

"Training not converging":

  • Adjust learning rate (try 0.01-0.05 range)
  • Increase training epochs
  • Check data quality and preprocessing

"Out of memory errors":

  • Reduce batch size for scEmbed
  • Process data in chunks
  • Use pre-tokenization for single-cell data

"StarSpace not found" (BEDspace):

For detailed troubleshooting and method-specific issues, consult the appropriate reference file.

© davila7, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 5 other files (references) in cli-tool/components/skills/scientific/geniml of davila7/claude-code-templates.

  • SKILL.md
  • references/bedspace.md
  • references/consensus_peaks.md
  • references/region2vec.md
  • references/scembed.md
  • references/utilities.md

Open the folder on GitHubat commit 46b4d8b

Used in 11 other repositories

We found 13 copies of this SKILL.md (exact, near-identical or edited) in other folders, from 11 other GitHub owners. This page covers the copy in davila7/claude-code-templates, which our catalogue first saw on October 7, 2026.

Compare with similar skills

Geniml next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

Geniml compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
Geniml this skilldavila7/claude-code-templates32k11 repos~2.5kAutomated safety check: PassMIT
Evo2JimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
ScgptJimLiu/science-skills2274 repos~1.3kAutomated safety check: PassApache-2.0
Celltype Specificity ProfilerClawBio/ClawBio1.2k—~4.3kAutomated safety check: PassMIT
Umap Tsne Analysisaipoch/medical-research-skills2k—~2.7kAutomated safety check: PassMIT
Genimlaipoch/medical-research-skills2k—~1.9kAutomated safety check: PassMIT

Similar skills

  • Evo2

    JimLiu/science-skills

    Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Scgpt

    JimLiu/science-skills

    Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology.

    227 GitHub starsUsed in 4 repos~1.3k tokens
    Research & ScienceAuto-check passed
  • Given a gene and a single-cell atlas, compute how cell-type-specific its expression is — the tau specificity index, Sarle's expression bimodality coefficient, and the cell types that drive the…

    1.2k GitHub stars~4.3k tokensUpdated today
    Research & ScienceAuto-check passed
  • Umap Tsne Analysis

    aipoch/medical-research-skills

    A skill your agent uses when performing sample-level dimensionality reduction and visualization on abundance or OTU-style matrices with a companion group file, generating UMAP and/or t-SNE…

    2k GitHub stars~2.7k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Geniml

    aipoch/medical-research-skills

    Machine learning toolkit for genomic interval (BED) data; use it when you need to tokenize BED collections and train embeddings for regions/cells/labels, build consensus peak universes, or run…

    2k GitHub stars~1.9k tokensUpdated 22 days ago
    Research & ScienceAuto-check passed
  • Sc Metacell

    TianGzlab/OmicsClaw

    Load when aggregating single cells into metacells (sample-aware coarse-grained pseudo-cells) on a normalised scRNA AnnData via SEACells or KMeans on a low-D embedding.

    161 GitHub stars~1.3k tokensUpdated yesterday
    Research & ScienceAuto-check passed

More from davila7/claude-code-templates

All 478 skills in this repo
  • Perplexity Web Search

    davila7/claude-code-templates

    Runs web-grounded searches through Perplexity's Sonar models over OpenRouter for current events, recent literature and cited facts beyond the model's training cutoff.

    32k GitHub starsUsed in 11 repos~3.5k tokens
    Auto-check: notes
  • Neuropixels Data Analysis

    davila7/claude-code-templates

    Analyzes Neuropixels recordings from SpikeGLX or Open Ephys through preprocessing, drift correction, Kilosort4 spike sorting, quality metrics and curation.

    32k GitHub starsUsed in 9 repos~2.8k tokens
    Auto-check passed
  • Scientific Venue Templates

    davila7/claude-code-templates

    Supplies LaTeX templates and formatting rules for journals, conferences, posters, and grant proposals, then can check a draft against them.

    32k GitHub starsUsed in 9 repos~5.1k tokens
    Auto-check: notes
  • Brand Voice Content Creator

    davila7/claude-code-templates

    Analyzes a brand's existing writing to lock in a consistent voice, then builds SEO blog posts and platform-specific social content around it.

    32k GitHub starsUsed in 3 repos~1.9k tokens
    Auto-check passed
  • CAPA Officer

    davila7/claude-code-templates

    Guides corrective and preventive action (CAPA) work in a quality management system, from initiation and root cause analysis through effectiveness verification.

    32k GitHub starsUsed in 1 repo~2k tokens
    Auto-check passed
  • Fda Consultant Specialist

    davila7/claude-code-templates

    Senior FDA consultant and specialist for medical device companies including HIPAA compliance and requirement management.

    32k GitHub starsUsed in 1 repo~2.7k tokens
    Auto-check passed

Questions about Geniml

What does Geniml do?

This skill should be used when working with genomic interval data (BED files) for machine learning tasks. Geniml is an agent skill from davila7/claude-code-templates. This skill should be used when working with genomic interval data (BED files) for machine learning tasks.

When should I use Geniml?

Geniml fits situations like: training region embeddings (Region2Vec; single-cell ATAC-seq analysis (scEmbed); building consensus peaks (universes); any ML-based analysis of genomic regions.

How do I install Geniml in Claude Code?

Run `npx skills add davila7/claude-code-templates --skill geniml -a claude-code`. Or copy the skill folder (cli-tool/components/skills/scientific/geniml in davila7/claude-code-templates) into .claude/skills/geniml in your project. Claude Code loads it when a task matches its description.

How do I install Geniml in Codex?

Run `npx skills add davila7/claude-code-templates --skill geniml -a codex`. Or copy the skill folder (cli-tool/components/skills/scientific/geniml in davila7/claude-code-templates) into .agents/skills/geniml in your project. Codex loads it when a task matches its description.

Can I use Geniml in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add davila7/claude-code-templates --skill geniml -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/geniml, .gemini/skills/geniml, .github/skills/geniml and .opencode/skills/geniml in your project.

What does Geniml need to run?

Going by SKILL.md and its folder, Geniml needs the command-line tools its instructions call (uv). Our summary lists: Python 3.

Does Geniml access the network?

SKILL.md names 2 domains. In commands or code: github.com; the agent is likely to contact it when it follows the instructions. As links in the text: docs.bedbase.org. This is read from the text; nothing was executed.

Is Geniml safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does Geniml use?

Geniml is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does Geniml use?

About 2.5k tokens (SKILL.md is roughly 10k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 6.7k tokens, read only when the agent opens those files.

What are the alternatives to Geniml?

Skills that share tags, products or a category with Geniml: Evo2 (JimLiu/science-skills, 227 stars), Scgpt (JimLiu/science-skills, 227 stars), Celltype Specificity Profiler (ClawBio/ClawBio, 1.2k stars) and Umap Tsne Analysis (aipoch/medical-research-skills, 2k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains Geniml?

davila7 (a GitHub user) maintains it in davila7/claude-code-templates, which has 32,483 GitHub stars. The repository holds 478 skills in this directory. The repository was last updated on October 9, 2026.

Source: davila7/claude-code-templates on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.