Agent skill

tangermeme Genomic Model Analysis

by jmschrei in jmschrei/tangermeme

Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.

MITAuto-check passedResearch & Science

Install tangermeme Genomic Model Analysis

skills CLI
$ npx skills add jmschrei/tangermeme --skill tangermeme -a claude-code

Project install by default; add -g for ~/.claude/skills/.

GitHub CLI
$ gh skill install jmschrei/tangermeme tangermeme --agent claude-code

Project scope by default; add --scope user for a personal install. Needs GitHub CLI 2.90.0 or later (public preview).

Manual copy
$ git clone --depth 1 https://github.com/jmschrei/tangermeme.git skills-src && mkdir -p .claude/skills && cp -r skills-src/tangermeme/_skills/data .claude/skills/tangermeme && rm -rf skills-src

Use ~/.claude/skills/ instead of .claude/skills for a personal install. The folder must contain SKILL.md.

Claude Code skills documentation · loads skills from .claude/skills/

Facts

Skill name
tangermeme
GitHub stars
316
Token cost
~1.6k tokens
SKILL.md length
632 words
Files
15 (incl. references)
Skills in repo
1
Repo updated
First seen
Licence
MIT

At a glance

Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design.

  • Computing DeepLIFT/SHAP attributions for a genomic deep learning model
  • SKILL.md covers Two cross-cutting concepts…, Task → reference file, Quick orientation (the rest of… and Conventions used throughout
  • Instructions only: no scripts, shell commands, URLs or credentials in SKILL.md
  • Scoring variant effects or running saturation mutagenesis on a sequence

What it does

tangermeme is a library for asking what a genomic sequence-to-function model learned once training is done. It offers atomic sequence operations, batched prediction, attribution, perturbation experiments and sequence design, and it stays assumption-free: any PyTorch model, any alphabet, raw outputs returned rather than distances. This skill is a router that sends the agent to one reference file per topic and insists they be read before any tangermeme code is written, since several functions have non-obvious defaults.

Two ideas are explained first. The func argument lets ablate, marginalize, space, variant_effect and product functions take any function with the model-and-inputs signature, so swapping predict for deep_lift_shap turns a predictions experiment into an attributions one. Models must also be wrapped so that calling them returns a single tensor, shaped batch by outputs for DeepLIFT/SHAP and pisa. A task table then points to references for a notebook walkthrough, attributions, saturation mutagenesis, model comparison, variant effects, seqlets, motif effects, loading loci and plotting.

When your agent uses it

  • Computing DeepLIFT/SHAP attributions for a genomic deep learning model
  • Scoring variant effects or running saturation mutagenesis on a sequence
  • Wrapping a multi-output PyTorch model so tangermeme can analyze it
  • Loading loci, FASTA, bigWig, MEME or VCF data for model analysis

Example prompts

  • “Compute DeepLIFT/SHAP attributions for my trained accessibility model on the peaks in peaks.bed.”
  • “Score the effect of every SNP in variants.vcf with my sequence-to-function model.”
  • “Run saturation mutagenesis on this enhancer sequence and plot the result.”
  • “Wrap my two-headed PyTorch model so tangermeme can compute attributions.”

Requirements

  • The tangermeme library
  • A trained PyTorch genomic model

What it can do on your machine

Read from SKILL.md and the folder at commit cebd9b1. It shows what the files ask for, not the result of running them.

  • Tool permissions

    Pre-approves nothing: there is no allowed-tools line, so your agent's usual permission prompts apply.

    From allowed-tools in the SKILL.md frontmatter.

  • Runs code

    No scripts in the folder and no shell commands in SKILL.md.

    From the folder's file list and the shell code blocks in SKILL.md.

  • Network

    No URLs in SKILL.md.

    From URLs in SKILL.md, links to its own repository left out.

  • Credentials

    Names no API keys, tokens, secrets or passwords.

    From names ending in _API_KEY, _TOKEN, _SECRET, _KEY or _PASSWORD in SKILL.md.

Context cost

tangermeme Genomic Model Analysis loads about 1.6k tokens when it runs, and up to ~26k if it reads all its reference files. Until then it costs about 136 tokens; SKILL.md has 632 words of instructions outside code blocks.

Always · name and description, kept in context so the agent knows when to use it
~136
When it runs · the whole SKILL.md, loaded when a task matches
~1.6k
With references · SKILL.md plus every file in references/, read only if the agent opens them
~26k

Estimates: characters ÷ 4, the usual rule of thumb; real counts depend on the model's tokenizer. Scripts and assets cost tokens only if the agent reads them.

Safety

Auto-check passed

The automated check found no risky patterns in SKILL.md.

Automated static check — not a guarantee. Review scripts before installing. It scans the text of SKILL.md for risky patterns (piping downloads into a shell, reading credential files, hidden Unicode, destructive commands); files beside SKILL.md are not scanned.

SKILL.md

The full file from jmschrei/tangermeme at commit cebd9b1, republished under its MIT licence (© jmschrei). 632 words, ~1,573 tokens.

Download SKILL.mdSave it as .claude/skills/tangermeme/SKILL.md (or your agent's skills folder). This skill also uses 14 other files; get the full folder from GitHub.
name
tangermeme
description
Use for any task involving the tangermeme library — post-training analysis of genomic sequence-to-function (S2F) deep learning models. Triggers on predictions, DeepLIFT/SHAP attributions, marginalization/ablation/spacing of motifs, saturation mutagenesis (ISM), variant effect scoring, sequence design, seqlet calling, loading loci/FASTA/bigWig/MEME/VCF, or wrapping a PyTorch genomic model for these analyses. This is a router skill — read the relevant file under references/ for details and footguns before writing tangermeme code.

tangermeme

tangermeme answers the "what did my genomic model learn, and what do I do with it after training" question. It provides atomic sequence operations, batched prediction, attribution, perturbation experiments, and sequence design — all deliberately assumption-free (any PyTorch model, any alphabet, raw outputs returned rather than distances).

This skill is a router. Each topic below has a detailed reference file with the exact signatures and footguns. Read the relevant reference file before writing code — do not rely on memory of the API, because several functions have non-obvious defaults (silent wrong-output selection, variable-length returns, reproducibility traps).

Two cross-cutting concepts (read these first if unsure)

  • The func= plug-point (references/func-pattern.md) — ablate, marginalize, space, variant_effect.*, and product.* (where it is the first positional argument) all accept func(model, X, args=, **kwargs). Swapping predict for deep_lift_shap turns a "predictions before/after" experiment into an "attributions before/after" one. Covers the additional_func_kwargs collision trap. This is what makes the library compose.

  • Wrapping models (references/model-wrapping.md) — tangermeme assumes y = model(X) returns a single tensor; deep_lift_shap and pisa need it to be (batch, n_outputs). Real multi-input / multi-output models must be wrapped first. Read this before attribution or design on any non-trivial model. Data preprocessing or output post-processing should be handled in custom wrappers rather than in custom functions.

Task → reference file

If the task is…Read
starting from scratch — set up a notebook to load a model and run predictions → attributions → seqlets → motif tests, end to endreferences/notebook-walkthrough.md
attribution via DeepLIFT/SHAP — "which bases drive this prediction", attribution logos, hypothetical contributions for CWMsreferences/deep_lift_shap.md
attribution via ISM / saturation mutagenesis — the forward-pass alternative; use it when DeepLIFT/SHAP convergence deltas are too high, an op can't be registered, or the model is massively multi-taskreferences/saturation_mutagenesis.md
comparing predictions/attributions across N models (replicates, architectures, ensembles)references/comparing-models.md
effect of a motif / region: marginalize, ablate, spacing between motifsreferences/motif-effects.md
scoring variant effects (substitution / deletion / insertion, from a VCF)references/variant-effect.md
calling seqlets from attributions (recursive / TF-MoDISco)references/seqlets.md
annotating / counting motifs — TOMTOM/FIMO labels, co-occurrence, spacingreferences/annotate.md
running a function over a product of inputs (sequence × cell-state × …)references/product.md
plotting logos and drawing seqlet/motif annotations on themreferences/plot.md
composing predict / deep_lift_shap / saturation_mutagenesis through a perturbation fnreferences/func-pattern.md
adapting a multi-input/output PyTorch model to the tangermeme contractreferences/model-wrapping.md
loading sequences/signals at loci, reading FASTA/bigWig/BED/MEME/VCFreferences/io-loci.md
designing sequences to hit a target output (screen / greedy / beam substitution)references/design.md
Show full SKILL.md (249 more words)Show less

Quick orientation (the rest of the library)

These are well-covered by the official tutorials and are mostly single-call ops — no dedicated reference file, but here is where to look:

  • tangermeme.predict.predict — batched, memory-efficient inference. Returns the model's parameter dtype (override with dtype=) and upcasts each batch from X, so int8 sequences go straight in; multi-output models return a list. Satisfies the func= contract.
  • tangermeme.ersatz — atomic sequence ops: insert, substitute, multisubstitute, delete, randomize, shuffle, dinucleotide_shuffle, local_dinucleotide_shuffle (most motif-add ops substitute, preserving length; start/end confine shuffles to a region). All but the two dinucleotide shuffles accept unknown characters in X as all-zero columns; dinucleotide_shuffle does with allow_N=True, shuffling each as a fifth character.
  • tangermeme.utils — one_hot_encode (returns int8), characters, random_one_hot, reverse_complement (single sequence), pwm_consensus, set_seed, gc_content, etc.
  • tangermeme.pisa.pisa — per-position (PISA) attribution reusing the DLS hooks. Footguns: some paths return tensors on the input device, not CPU; and it does not upcast X, so pass float — it is the one entry point int8 fails on.
  • tangermeme.kmers — k-mer counts, batched k-mers, gapped k-mers.

(Motif scanning — FIMO/TOMTOM — is not in tangermeme; it lives in the external memelite / memesuite-lite package. annotate_seqlets wraps TOMTOM internally.)

Conventions used throughout

  • Tensor layout: (batch, channels, length); channels = one-hot alphabet axis (default ['A','C','G','T']).
  • Variable naming: X/y observed, X_bar/y_bar designed/target, y_hat predictions, X_attr attributions.
  • Device defaults: predict, deep_lift_shap, pisa, design.*, saturation_mutagenesis, product.* take device=None → CUDA if available else CPU; the model's original device + training mode are restored afterward.
  • Perturbation functions return NamedTuples (PerturbationResult, etc.) — unpack positionally or by attribute; isinstance(result, tuple) is True.

© jmschrei, MIT. Rendered from Markdown: HTML in the file is shown as text, images as links, and headings moved down two levels. Raw file

Files

SKILL.md and 14 other files (references) in tangermeme/_skills/data of jmschrei/tangermeme.

  • SKILL.md
  • references/annotate.md
  • references/comparing-models.md
  • references/deep_lift_shap.md
  • references/design.md
  • references/func-pattern.md
  • references/io-loci.md
  • references/model-wrapping.md
  • references/motif-effects.md
  • references/notebook-walkthrough.md
  • references/plot.md
  • references/product.md
  • references/saturation_mutagenesis.md
  • references/seqlets.md
  • references/variant-effect.md

Open the folder on GitHubat commit cebd9b1

Compare with similar skills

tangermeme Genomic Model Analysis next to the 5 skills that share the most tags, products or categories with it. Stars are the repository's; “used in” counts other GitHub owners with a copy.

tangermeme Genomic Model Analysis compared with similar skills
SkillStarsUsed inTokensAuto-checkLicenceRepo updated
tangermeme Genomic Model Analysis this skilljmschrei/tangermeme316—~1.6kAutomated safety check: PassMIT
Cellxgene Censusdavila7/claude-code-templates32k11 repos~3.8kAutomated safety check: PassMIT
Pixi Environment Builderxuzhougeng/wisp-science1k—~3.7kAutomated safety check: PassAGPL-3.0
Alphagenome Predictionsgenomicsxai/alphagenome-pytorch162—~868Automated safety check: PassApache-2.0
PyTorch Lightning TrainingOrchestra-Research/AI-Research-SKILLs13k6 repos~2.3kAutomated safety check: PassMIT
Ray Train Distributed TrainingOrchestra-Research/AI-Research-SKILLs13k2 repos~2.7kAutomated safety check: PassMIT

Similar skills

  • Cellxgene Census

    davila7/claude-code-templates

    Query CZ CELLxGENE Census (61M+ cells). An agent skill from davila7/claude-code-templates.

    32k GitHub starsUsed in 11 repos~3.8k tokens
    Research & ScienceAuto-check passed
  • Pixi Environment Builder

    xuzhougeng/wisp-science

    A skill your agent uses when creating, migrating, or debugging pixi environments, especially for scientific Python, bioinformatics, single-cell analysis, CUDA/PyTorch, Jupyter/VS Code kernels…

    1k GitHub stars~3.7k tokensUpdated today
    Research & ScienceAuto-check passed
  • Alphagenome Predictions

    genomicsxai/alphagenome-pytorch

    Run AlphaGenome-PyTorch to get genomic track predictions — via the agt predict CLI (single locus, BED regions, whole chromosomes, raw FASTA sequences, or per-gene count tables/AnnData), variant…

    162 GitHub stars~868 tokensUpdated 24 days ago
    AI & LLM EngineeringAuto-check passed
  • PyTorch Lightning Training

    Orchestra-Research/AI-Research-SKILLs

    Shows how to organize PyTorch training with Lightning's LightningModule and Trainer, covering validation, DDP, callbacks and learning-rate scheduling.

    13k GitHub starsUsed in 6 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed
  • Ray Train Distributed Training

    Orchestra-Research/AI-Research-SKILLs

    Scales PyTorch, TensorFlow and Hugging Face training from a single GPU to multi-node clusters with Ray Train, including Ray Tune sweeps and checkpoint recovery.

    13k GitHub starsUsed in 2 repos~2.7k tokens
    AI & LLM EngineeringAuto-check passed
  • ML Engineer

    davila7/claude-code-templates

    Build production ML systems with PyTorch 2.x, TensorFlow, and modern ML frameworks.

    32k GitHub starsUsed in 9 repos~2.3k tokens
    AI & LLM EngineeringAuto-check passed

Works with

Questions about tangermeme Genomic Model Analysis

What does tangermeme Genomic Model Analysis do?

Routes agents to the right tangermeme reference for analyzing trained genomic deep learning models, from attributions and motif experiments to variant effects and design. tangermeme is a library for asking what a genomic sequence-to-function model learned once training is done. It offers atomic sequence operations, batched prediction, attribution, perturbation experiments and sequence design, and it stays assumption-free: any PyTorch model, any alphabet, raw outputs returned rather than distances.

When should I use tangermeme Genomic Model Analysis?

tangermeme Genomic Model Analysis fits situations like: computing DeepLIFT/SHAP attributions for a genomic deep learning model; scoring variant effects or running saturation mutagenesis on a sequence; wrapping a multi-output PyTorch model so tangermeme can analyze it; loading loci, FASTA, bigWig, MEME or VCF data for model analysis.

How do I install tangermeme Genomic Model Analysis in Claude Code?

Run `npx skills add jmschrei/tangermeme --skill tangermeme -a claude-code`. Or copy the skill folder (tangermeme/_skills/data in jmschrei/tangermeme) into .claude/skills/tangermeme in your project. Claude Code loads it when a task matches its description.

How do I install tangermeme Genomic Model Analysis in Codex?

Run `npx skills add jmschrei/tangermeme --skill tangermeme -a codex`. Or copy the skill folder (tangermeme/_skills/data in jmschrei/tangermeme) into .agents/skills/tangermeme in your project. Codex loads it when a task matches its description.

Can I use tangermeme Genomic Model Analysis in Cursor, Gemini CLI or GitHub Copilot?

Cursor, Gemini CLI, GitHub Copilot and OpenCode also load SKILL.md folders. With the skills CLI, run `npx skills add jmschrei/tangermeme --skill tangermeme -a cursor` (or -a gemini-cli, github-copilot or opencode for the others). To copy it by hand, put the folder in .cursor/skills/tangermeme, .gemini/skills/tangermeme, .github/skills/tangermeme and .opencode/skills/tangermeme in your project.

What does tangermeme Genomic Model Analysis need to run?

SKILL.md names no scripts, command-line tools or credentials: tangermeme Genomic Model Analysis is instructions for the agent only. Our summary lists: The tangermeme library; A trained PyTorch genomic model.

Does tangermeme Genomic Model Analysis access the network?

SKILL.md contains no URLs. Any network use would come from the scripts or tools the agent runs. This is read from the text; nothing was executed.

Is tangermeme Genomic Model Analysis safe to install?

Our automated static check of SKILL.md found no risky patterns, such as piping downloads into a shell, reading credential files or hidden Unicode. It is not a guarantee. Review the folder before installing.

What licence does tangermeme Genomic Model Analysis use?

tangermeme Genomic Model Analysis is published under the MIT licence (the repository's licence). It allows redistribution, so the full SKILL.md is shown on this page.

How many tokens does tangermeme Genomic Model Analysis use?

About 1.6k tokens (SKILL.md is roughly 6.3k characters). Agents keep only the skill's name and description in context until a task matches; then they load SKILL.md in full. Its references folder adds about 25k tokens, read only when the agent opens those files.

What are the alternatives to tangermeme Genomic Model Analysis?

Skills that share tags, products or a category with tangermeme Genomic Model Analysis: Cellxgene Census (davila7/claude-code-templates, 32k stars), Pixi Environment Builder (xuzhougeng/wisp-science, 1k stars), Alphagenome Predictions (genomicsxai/alphagenome-pytorch, 162 stars) and PyTorch Lightning Training (Orchestra-Research/AI-Research-SKILLs, 13k stars). The comparison table on this page puts their stars, adoption, token cost, safety result and licence side by side.

Who maintains tangermeme Genomic Model Analysis?

jmschrei (a GitHub user) maintains it in jmschrei/tangermeme, which has 316 GitHub stars. The repository was last updated on October 7, 2026.

Source: jmschrei/tangermeme on GitHub. Facts on this page come from the repository at the commit we read; the author's words are quoted as theirs.